Plug-in large model access system for embedded equipment

By using a pluggable large model access system, the problems of high cost of multi-model adaptation, large resource consumption and low communication efficiency when embedded devices access large cloud models are solved, and efficient and secure cloud interaction is achieved in resource-constrained environments.

CN120980144APending Publication Date: 2025-11-18ZHUHAI HUGE IC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510995168.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

When embedded devices connect to large cloud models, there are problems such as high cost of multi-model adaptation, excessive resource consumption, and insufficient communication efficiency and security.

Method used

The system adopts a pluggable large model access architecture, including a session layer, a platform layer, and a transport layer. It provides standard calling interfaces, various large model plug-ins, and communication protocols, and dynamically selects the optimal communication protocol to achieve unified access to different cloud platforms and large models.

Benefits of technology

It reduces platform switching and integration development costs, improves transmission efficiency and security, and enables efficient interaction of resource-constrained embedded devices in weak network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980144A_ABST
    Figure CN120980144A_ABST
Patent Text Reader

Abstract

The invention discloses a plug-in large model access system for embedded equipment, and relates to the field of large model application. The system comprises a session layer used for providing a standard calling interface and defining different large model calling modes, and selecting a target large model calling mode based on a service type to send a service to a platform layer; the platform layer comprises a plurality of large model plug-ins of different types, and is used for serializing a protocol and a format of a service by adopting a target large model plug-in to obtain a calling request matched with the protocol and the format of the target large model, and sending the calling request to the transmission layer; the deserialization analysis layer is used for carrying out deserialization analysis on the calling request result and sending a deserialization analysis result to the session layer; and the transmission layer comprises a plurality of communication protocols and is used for determining a target communication protocol according to a network environment and service requirements, sending a calling request to a target large model and sending a calling request result to the platform layer. According to the invention, the technical problems of high multi-model adaptation cost and low communication efficiency are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model application technology, and in particular to a pluggable large model access system for embedded devices. Background Technology

[0002] With the deep integration of edge computing and artificial intelligence, the demand for embedded devices to access large AI models is growing. Typical application scenarios include predictive maintenance of industrial IoT devices, multimodal interaction in smart homes, and intelligent analysis of health data from wearable devices. However, embedded devices generally suffer from resource constraints, making it difficult to deploy large models locally. Therefore, more and more smart terminal devices need to access cloud-based AI services via the network to complete intelligent tasks. However, this approach currently faces three major technical challenges: high cost of multi-model adaptation, as different cloud-based large models (such as OpenAIChatGPT, Baidu Wenxin Yiyan, and Alibaba Tongyi Qianwen) have significantly different API interfaces, authentication methods, and data formats, requiring the redevelopment of underlying communication modules for each adaptation; secondly, existing middleware does not optimize memory management and computing power scheduling for embedded systems, resulting in high ROM and RAM usage. Small embedded devices typically have ≤256KB of RAM, while mainstream AI access SDKs such as AWS IoTDevice SDK require >512KB, and Alibaba Cloud AIoT C-SDK alone consumes 300KB of memory for its TLS stack. Finally, it is difficult to balance functions such as long connection management, data compression, and security encryption in resource-constrained environments.

[0003] In existing technologies, the AWS Lambda + API Gateway solution uses AWS Lambda as the backend serverless computing engine, combined with the API Gateway to manage and invoke RESTful interfaces. Small embedded devices upload data to the API Gateway via HTTP / HTTPS protocols, which then forwards the requests to Lambda functions for processing. These Lambda functions can further interact with large cloud models (such as triggering inference tasks or retrieving results). The entire process requires no server resources and has excellent auto-scaling capabilities. However, when addressing the specific needs of small embedded devices accessing large cloud models, it has high requirements for the network environment, relies on public network access, and is difficult to adapt to devices accessing weak edge networks or private networks. Compared to embedded-specific SDKs that support low-power, lightweight communication, the API Gateway only supports the HTTP(S) protocol and cannot directly support protocols more suitable for embedded devices, such as MQTT and CoAP, adding an intermediate conversion layer and increasing system complexity and energy consumption.

[0004] In other technologies, the Alibaba Cloud AIoT C-SDK solution is adopted. This solution is based on the Alibaba Cloud AIoT C-SDK, a software toolkit specifically designed for embedded devices. It supports multiple lightweight communication protocols (such as MQTT, HTTP, and CoAP) and integrates core functions such as security authentication, OTA upgrades, and message subscription and publishing. Devices can directly connect to the Alibaba Cloud IoT platform through this SDK and, with the help of the platform's edge computing modules or function computing interfaces, achieve efficient interaction with large cloud models. For example, after the device uploads raw data, it is preprocessed by edge nodes, then triggers a large cloud model inference task, and finally the results are fed back to the device. However, compared to general cloud services (such as AWS Lambda), this SDK primarily serves the Alibaba Cloud ecosystem and lacks good compatibility with other cloud platforms, limiting consistent deployment in multi-cloud environments. Compared to the fully customizable FaaS architecture, the functional modules provided by the C-SDK are relatively fixed. If complex business logic needs to be implemented, users need to develop many additional intermediate components, resulting in poor flexibility.

[0005] Therefore, it is necessary to address the problems existing in current technologies, such as high multi-model adaptation costs, excessive resource consumption, and insufficient communication efficiency and security, when embedded devices access large cloud models. Summary of the Invention

[0006] This invention provides a pluggable large model access system for embedded devices, which solves the technical problems of high multi-model adaptation costs, excessive resource consumption, and insufficient communication efficiency and security in existing technologies when embedded devices access large cloud models. The technical solution is as follows:

[0007] In a first aspect, embodiments of the present invention provide a pluggable large-scale model access system for embedded devices, comprising:

[0008] The session layer is used to provide standard calling interfaces and define different large model calling methods, and selects the target large model calling method based on the business type to send the business to the platform layer;

[0009] The platform layer includes multiple large model plugins of different types. The platform layer is used to serialize the protocol and format of the business using the target large model plugin corresponding to the target large model calling method, to obtain a calling request that matches the target large model protocol and format, and to send the calling request to the transport layer; and to deserialize and parse the calling request result, and send the deserialization and parsing result to the session layer.

[0010] The transport layer includes multiple communication protocols, which are used to determine the target communication protocol based on the network environment and business requirements, and send the call request to the target large model through the target communication protocol, and send the call request result fed back from the large model to the platform layer.

[0011] In some embodiments of the present invention, the session layer is also used to feed back the deserialization result to the application layer.

[0012] In some embodiments of the present invention, the session layer is also used to record the state machine of the service, save and back up the service based on the target large model calling method, and perform breakpoint resume and / or retransmission of the service.

[0013] In some embodiments of the present invention, the large model invocation method includes at least automatic speech recognition, natural language generation, computer vision, natural language understanding, and intelligent agents.

[0014] In some embodiments of the present invention, the service includes metadata, which includes at least timestamp, device ID, model version number, priority tag and session ID.

[0015] In some embodiments of the present invention, the plurality of different types of large model plugins include at least Alibaba Cloud plugin, Tencent Cloud plugin, Volcano Bean Bun plugin, and iFlytek plugin.

[0016] In some embodiments of the present invention, the platform layer is further configured to perform segmentation processing of the service based on the data volume of the call request, and to perform package assembly operation on the call request results based on the dispersion of the call request results.

[0017] In some embodiments of the present invention, the communication protocol includes at least one of HTTP(S), WebSocket(S), MQTT, and CoAP.

[0018] In some embodiments of the present invention, the transport layer is further configured to acquire communication anomalies and define retry strategies based on the anomalies.

[0019] In some embodiments of the present invention, the transport layer is also used to establish long connections, short connections, and / or heartbeat keep-alive mechanisms.

[0020] The beneficial effects of the technical solutions provided by some embodiments of the present invention include at least the following: receiving call requests with a unified transmission format through a unified standard call interface API provided by the session layer avoids the problem of high adaptation costs caused by different authentication methods for different call requests, avoids underlying differences, defines the large model call method corresponding to different call requests when transmitting call requests based on the session layer, and then encapsulates various large model plugins through the platform layer, selects different target large model plugins based on different call requests, automatically serializes business data into the protocol format of the target large model, and adds platform support without modifying the core code; it realizes the embedded application of various large models; finally, based on the transport layer, it dynamically selects the optimal communication protocol according to the network environment (weak network / public network) and business requirements (real-time / security), thereby improving transmission efficiency. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 An exemplary system architecture diagram provided for this invention;

[0023] Figure 2 This is a system architecture diagram of an embodiment of the pluggable large model access system for embedded devices provided by the present invention;

[0024] Figure 3 This is a schematic diagram of metadata transmission in this invention;

[0025] Figure 4 This is a schematic diagram of the application process in this invention;

[0026] Figure 5 This is a schematic diagram of the core access system process in this invention;

[0027] Figure 6 This is a schematic diagram of the terminal device. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] It should be noted that the pluggable large model access system for embedded devices provided in this application is generally executed by the terminal device.

[0030] Figure 1An exemplary system architecture for a pluggable large-scale model access system for embedded devices, which can be applied to this application, is shown.

[0031] like Figure 1 As shown, the system architecture may include: terminal device 101 and server 102. Terminal device 101 and server 102 can communicate via a network, which serves as the medium for providing communication links between the various units. The network may include various types of wired or wireless communication links, such as: wired communication links including fiber optic cables, twisted-pair cables, or coaxial cables; and wireless communication links including Bluetooth communication links, Wi-Fi communication links, or microwave communication links.

[0032] It should be noted that the terminal device 101 and the server 102 can be either hardware or software. When the terminal device 101 and the server 102 are hardware, they can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the terminal device 101 and the server 102 are software, they can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module; no specific limitations are made here.

[0033] The terminal device of this application can be equipped with various communication client applications, such as video recording applications, video playback applications, voice interaction applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0034] A terminal device can be either hardware or software. When the terminal device is hardware, it can be various terminal devices with a display screen, including but not limited to smartphones, tablets, laptops, and desktop computers. When the terminal device is software, it can be installed on the terminal devices listed above. It can be implemented as multiple software programs or software modules (e.g., used to provide distributed services) or as a single software program or software module; no specific limitation is made here.

[0035] When the terminal device is hardware, it can also be equipped with a display device and a camera. The display device can be any device capable of displaying information, and the camera is used to capture video streams. For example, the display device can be a cathode ray tube display (CR), a light-emitting diode display (LED), an e-ink screen, a liquid crystal display (LCD), or a plasma display panel (PDP). Users can use the display device on the terminal device to view displayed text, images, videos, and other information.

[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is for illustrative purposes only. Depending on implementation needs, there can be any number of terminal devices, networks, and servers.

[0037] The following will be combined with the appendix Figure 2 This application provides a detailed description of the pluggable large model access system 20 for embedded devices provided in the embodiments of this application. Please refer to... Figure 2 This application provides a system architecture diagram for a pluggable large-scale model access system for embedded devices. For example... Figure 2 As shown, the system described in this application embodiment may include:

[0038] Session layer 21 is used to provide standard calling interfaces and define different large model calling methods, and select the target large model calling method based on the business type to send the business to the platform layer;

[0039] It should be noted that the session layer, as the portal to the access system, provides standard call interfaces to avoid underlying differences and solve the problem of high difficulty and low efficiency in embedding large models due to the need to adapt to different authentication methods for different large model business requirements.

[0040] It should be noted that the unified standard API is a custom interface, and no restrictions are imposed here. In a specific embodiment, the interface is session_send().

[0041] It should be noted that the large model invocation method is a technical framework that integrates modules such as Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), Computer Vision (CV), Natural Language Generation (NLG), and Agent to achieve collaborative processing of multimodal inputs and intelligent output generation. The large model invocation method includes, but is not limited to, ASR, NLU, CV, NLG, and Agent. Large model plugins include at least Alibaba Cloud, Tencent Cloud, Huoshan Doubao, and iFlytek plugins. Different large model invocation methods are processed through different or the same large model plugins. When different business needs, such as speech recognition or text recognition, can be processed simultaneously using Tencent Cloud or Alibaba Cloud, the target large model is selected based on a pre-defined priority. Therefore, through the unified standard calling interface provided by the session layer and different large model invocation methods, the application layer only needs to focus on "what model to call (WHAT)" rather than "how to call the model (HOW)".

[0042] Platform layer 22 includes multiple large model plugins of different types. The platform layer is used to serialize the protocol and format of the business using the target large model plugin corresponding to the target large model calling method, to obtain a calling request that matches the target large model protocol and format, and to send the calling request to the transport layer; and is used to deserialize and parse the calling request result, and send the deserialization and parsing result to the session layer.

[0043] It should be noted that by encapsulating the large models in the platform layer as large model plugins, and calling the corresponding plugins to process according to business needs, the goal of handling business without modifying the code and changing different large models is achieved.

[0044] It should be noted that the type, content, and generation method of the large model plugin are not limited here. The large model plugin is used to convert the general instructions for calling large models into the specific protocol of the target large model cloud platform, realize platform-specific authentication, dynamically generate timestamp+sign signatures, fill them into the HTTP header, and handle the platform-specific data fragmentation / compression logic, fragmenting into 128KB chunks and adding chunk_id tags.

[0045] Transport layer 23 includes multiple communication protocols, used to determine the target communication protocol according to the network environment and business requirements, and send the call request to the target large model through the target communication protocol, and send the call request result fed back from the large model to the platform layer.

[0046] It should be noted that the communication protocols include, but are not limited to, HTTP(S), WebSocket(S), MQTT, CoAP, etc., and the transport layer is an open protocol module. Furthermore, the most suitable protocol can be dynamically selected based on the network environment and business needs, thereby improving the efficiency and security of network transmission.

[0047] In some preferred embodiments, the session layer is also used to feed back the deserialization result to the application layer.

[0048] It should be noted that after obtaining the call results based on the large model, the format and protocol of the call results are different. Therefore, the large model plugin deserializes the call results, extracts key information from the deserialized data, and feeds it back to the application layer through a unified standard call interface API.

[0049] In some preferred embodiments, the session layer is also used to record the state machine of the service, save and back up the service based on the target large model calling method, and perform breakpoint resume and / or retransmission of the service.

[0050] It should be noted that the session is initialized first: a session instance is created based on the request parameters, and session-related parameters are configured. When the session is interrupted, the current request is canceled, and related resources (such as memory and threads) are released after the session ends.

[0051] Furthermore, the system automatically switches states based on event-triggered state changes and periodically polls the state to prevent deadlocks or blocking.

[0052] Furthermore, to ensure the reliability of data transmission and reception, asynchronous transmission and reception are supported, the sent / received data streams are cached, and breakpoint resumption is supported. Automatic handling of network anomalies such as data retransmission, packet loss detection, and out-of-order reassembly is implemented. A disconnection reconnection mechanism is also supported, automatically attempting to resume the session after a connection interruption.

[0053] In some preferred embodiments of the present invention, the service includes metadata, which includes at least timestamp, device ID, model version number, priority tag and session ID.

[0054] It should be noted that, as Figure 3As shown, metadata enables rapid identification of problem chains. It is passed from the application layer to session_send(), and the session layer supplements the key fields (session_id / timestamp). It travels with business data through the platform layer and transport layer to the cloud big model. The metadata returned by the cloud big model is used to optimize local decision-making. This design makes metadata the core information carrier that runs through the layers. It is the key to realizing lightweight intelligent decision-making. That is, in the resource-constrained embedded environment, it uses a very small storage cost (≈50 bytes / packet) to exchange for a qualitative improvement in operation and maintenance efficiency, security and observability.

[0055] In some embodiments of the present invention, the platform layer is further configured to perform segmentation processing of the service based on the data volume of the call request, and to perform package assembly operation on the call request results based on the dispersion of the call request results.

[0056] In this embodiment, when the amount of data transmitted in a single transmission is large, it is automatically fragmented and fragment identifiers are added. The complete data packet is then reassembled in order at the receiving end to ensure data integrity.

[0057] In some embodiments of this invention, the transport layer is used to establish, maintain, and disconnect communication connections, supporting long connections, short connections, and heartbeat keep-alive mechanisms. It also attempts to restore connections after network fluctuations or disconnections, achieving automatic reconnection. Furthermore, it provides a unified data sending and receiving interface, supports synchronous and asynchronous send / receive modes, supports QoS level control, and is suitable for business scenarios with varying reliability requirements. It automatically fragments large data volumes to improve transmission efficiency. It supports data compression algorithms (such as GZIP, LZ4, etc.) to reduce bandwidth consumption. It automatically captures communication anomalies (such as timeouts, disconnections, verification failures, etc.). It provides detailed error codes and log information for easy troubleshooting. It supports custom retry strategies (such as exponential backoff, fixed intervals, etc.). It supports TLS / SSL encrypted communication to ensure data transmission security. It provides functions such as identity authentication, access control, and signature verification.

[0058] In one specific embodiment of the present invention, please refer to Figure 4 and Figure 5 :

[0059] S1: Performs global initialization (global_init) for different systems and embedded platforms, including initializing system resources (memory pool, thread pool, event queue, etc.); registering OS abstract interfaces (malloc / free / memcpy, etc.); and configuring a list of supported models, supporting combined model calls (such as automatic speech recognition + natural language understanding + natural language generation). Only the models used by the business need to be registered, without compiling all of them, saving memory resources.

[0060] S2: Creates one or more sessions (session_init), each bound to a model. Configures the model name, session layer parameters (such as asynchronous send / receive queue size, receive buffer size, etc.), platform layer parameters (such as server-supported uplink and downlink audio data formats, setting timbre, volume, speech rate, etc.), and transfer layer parameters (such as timeout, retransmission count, etc.). These settings allow for finer-grained application control; otherwise, default parameters are provided. Simultaneously, it calls the platform object's operation set's model_init to initialize the platform layer and the transfer object's operation set's protocol_init to initialize the transport layer. Returns a session handle (session_handle) for subsequent calls.

[0061] S3: The system starts up, switches from the IDLE state to the CONNECTING state, and begins authentication and connection through the `model_connect` method of the platform object operation set. If it fails, it remains in the CONNECTING state, repeatedly attempts, and reports the status or reason for failure (such as communication error, server return of `err_msg`, etc.). If successful, it switches from the CONNECTING state to the CONNECTED state, and the middleware process begins to continuously listen to and process the data transmission and reception queues.

[0062] S4: The application begins asynchronous data sending and receiving. This process is non-blocking and polling. The application ensures real-time coverage of both the sending and receiving processes. The specific steps are as follows:

[0063] S4-1: Sending process:

[0064] S4-1-1: The application collects data app_data and calls session_send to send it. If the circular buffer (tx_rb, first-in-first-out) of tx is not full, app_data is copied and added to tx_rb with a specific data structure (such as containing source data, data length, offset, etc.). If it is full, it returns again to inform the application.

[0065] S4-1-2: When the system process detects that there is data in tx_rb, it extracts the bottom data and marks it as unsend data. It checks the model name and type corresponding to the data. If the check passes, it calls the model_send method of the platform object operation set to send the data.

[0066] S4-1-3: Data is encapsulated, encrypted (e.g., MD5), compressed (e.g., base64), and packetized (if the data packet is large) in the platform layer. After completion, it is sent through the protocol_send of the transfer object operation set.

[0067] S4-1-4: Data is processed in the transport layer by communication protocols (such as HTTP(S), WebSocket(S), MQTT, CoAP, etc.), which also involves SSL / TLS encryption. Finally, it is sent out by the underlying protocol stack.

[0068] S4-1-5: If the transmission fails, an exception code is returned layer by layer, triggering a retransmission mechanism. Once the retransmission limit is reached, the packet is discarded and reported to the application layer. If transmission is successful, the actual amount of data sent is returned. If the actual amount of data sent equals the expected amount, it indicates that the packet transmission is complete, space is released, and a success code is returned. If not, it indicates that the packet was not fully transmitted (only a portion was sent), the offset in the previously reserved data information is updated, and `again` is returned. Upon receiving this, the session layer pauses retrieving packets from `tx_rb` and continues sending the incomplete data packets.

[0069] S4-2: Receiving Procedure:

[0070] S4-2-1: The system process continuously calls the `model_recv` method of the `platform` object operation set to listen for data. If no data is received, it waits for the next call. If data is received, it is first received by the `protocol_recv` method of the `transfer` object operation set, and the communication protocol is parsed. The data integrity is verified, and the data is assembled into a complete packet and sent to the platform layer.

[0071] S4-2-2: After receiving the data, the platform layer verifies the data integrity, performs deserialization and parsing of the data, including de-encapsulation, decryption, and decompression of the platform protocol, and extracts key information to feed back to the session layer.

[0072] S4-2-3: The session layer checks the status of the circular buffer (rx_rb, first-in-first-out) of rx. If rx_rb is not full, it adds data in; if it is full, it waits for the application to call session_recv to retrieve the data.

[0073] S5: In abnormal situations, such as communication failures or server requests for device disconnection, the CONNECTED state will switch to the DISCONNECT state, triggering resource reclamation and the disconnection / reconnection process. For other "abnormalities," such as the application actively interrupting the transmission, the CONNECTED state will switch to the RECYCLE state, triggering only resource reclamation. If resource reclamation fails, the reconnection mechanism will be triggered directly, switching from the RECYCLE state to the DISCONNECT state and initiating the disconnection / reconnection process.

[0074] S6: The system exits, the application calls session_deinit to destroy the session, and finally calls global_deinit to reclaim global resources, and the process ends.

[0075] This invention offers the following advantages: It receives call requests with a unified transmission format via a unified standard API provided by the session layer, thus avoiding the high adaptation costs caused by different authentication methods for different call requests and avoiding underlying differences. When transmitting call requests based on the session layer, it defines the large model call method corresponding to different call requests. Subsequently, it encapsulates various large model plugins through the platform layer, selecting different target large model plugins based on different call requests, and automatically serializing business data into the protocol format of the target large model. Adding platform support requires no modification to the core code; it enables embedded applications of various large models; finally, based on the transport layer, it dynamically selects the optimal communication protocol according to the network environment (weak network / public network) and business requirements (real-time / security), improving transmission efficiency. Therefore, this application proposes a modular, pluggable, layered architecture-based communication intermediate access system. Through standardized interface design, asynchronous non-blocking polling, and dynamic resource management, it achieves unified access to different cloud platforms and large model services, significantly reducing the development costs of platform switching and integration. It also supports efficient operation on resource-constrained embedded devices, improving the interaction efficiency between devices and cloud-based large models while ensuring data transmission security.

[0076] Please see Figure 6 This document provides a schematic diagram of the structure of a terminal device according to an embodiment of this application. Figure 6 As shown, the terminal device 600 may include: at least one processor 601, at least one network interface 604, user interface 603, memory 605, and at least one communication bus 602.

[0077] The communication bus 602 is used to enable communication between these components.

[0078] The user interface 603 may include a display screen and a camera. Optionally, the user interface 603 may also include a standard wired interface and a wireless interface.

[0079] The network interface 604 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0080] The processor 601 may include one or more processing cores. The processor 601 connects to various parts within the terminal device 600 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 605, and by calling data stored in the memory 605. Optionally, the processor 601 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 601 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 601 and may be implemented as a separate chip.

[0081] The memory 605 may include random access memory (RAM) or read-only memory. Optionally, the memory 605 may include a non-transitory computer-readable storage medium. The memory 605 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 605 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 605 may also be at least one storage device located remotely from the aforementioned processor 601. Figure 6 As shown, the memory 605, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and application programs.

[0082] exist Figure 6In the terminal device 600 shown, the user interface 603 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 601 can be used to call the application program stored in the memory 605.

[0083] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0084] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. A pluggable large-scale model access system for embedded devices, characterized in that, include: The session layer is used to provide standard calling interfaces and define different large model calling methods, and selects the target large model calling method based on the business type to send the business to the platform layer; The platform layer includes multiple large model plugins of different types. The platform layer is used to serialize the protocol and format of the business using the target large model plugin corresponding to the target large model calling method, to obtain a calling request that matches the target large model protocol and format, and to send the calling request to the transport layer. And it is used to deserialize and parse the results of the call request, and send the deserialization and parsing results to the session layer; The transport layer includes multiple communication protocols, which are used to determine the target communication protocol based on the network environment and business requirements, and send the call request to the target large model through the target communication protocol, and send the call request result fed back from the large model to the platform layer.

2. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The session layer is also used to feed back the deserialization result to the application layer.

3. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The session layer is also used to record the state machine of the service, save and back up the service based on the target large model calling method, and perform breakpoint resume and / or retransmission of the service.

4. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The large model invocation methods include at least automatic speech recognition, natural language generation, computer vision, natural language understanding, and intelligent agents.

5. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The service includes metadata, which includes at least timestamps, device IDs, model version numbers, priority tags, and session IDs.

6. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The various types of large model plugins include at least the Alibaba Cloud plugin, Tencent Cloud plugin, Huoshan Doubao plugin, and iFlytek plugin.

7. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The platform layer is also used to perform business segmentation processing based on the data volume of the call request, and to perform package assembly operation on the call request results based on the dispersion of the call request results.

8. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The communication protocol includes at least one of HTTP(S), WebSocket(S), MQTT, and CoAP.

9. The pluggable large-scale model access system for embedded devices according to claim 1, characterized in that, The transport layer is also used to acquire communication anomalies and define retry strategies based on the anomalies.

10. The pluggable large model access system for embedded devices according to claim 1, characterized in that, The transport layer is also used to establish long connections, short links, and / or heartbeat keep-alive mechanisms.

Citation Information

Patent Citations

  • Multi-source large model service standardized management and control system and method

    CN119512620A

  • JDK8-based heterogeneous service large model dynamic adaptation and intelligent scheduling method and system

    CN120281824A