Task processing method and device, equipment and computer medium
By using a unified base model architecture and dynamic parameter issuance mechanism on the client, the multi-service AI capabilities are quickly deployed, and the problems of resource waste and redundancy in the existing technology are solved, and efficient AI processing is achieved.
Patent Information
- Application Number
- CN202510468257.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
AI Technical Summary
In client scenarios, existing ways to solve AI needs are wasteful, including the use of online services and local artificial algorithms, resulting in waste of resources and redundancy.
Through the unified base model architecture combined with the dynamic parameter issuance mechanism, a system with rapid deployment of multi-service AI capabilities on mobile/embedded devices can be realized. The specific method includes determining task information in response to a request to detect a task execution, obtaining model parameter information corresponding to a task category, applying these parameters to the target base model to generate a model for performing the task, and processing the to-be-processed data based on the model.
When performing tasks, this solution simplifies the model determination process, saves model layout costs, and reduces storage and computing resources when solving AI requirements based on local models.
Smart Images

Figure CN119988040A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a task processing method, device, equipment and computer medium. Background Art
[0002] At present, many businesses in some clients have used or plan to use the in-end offline AI (Artificial Intelligence) reasoning capabilities to solve specific engineering problems in the end. At present, some clients mainly use two methods to solve AI needs. One is to use online services (such as enterprise Q&A, meeting minutes, screenshot OCR) to solve AI needs, and the other is to use artificial algorithms, such as some AI conclusions cached locally to solve AI needs. Among them, when using online services, many daily specific problems do not require large-scale, high-cost reasoning models, which will lead to waste of resources. In addition, when using artificial algorithms, multiple business scenarios in the client will independently deploy multiple models (such as phishing detection / network diagnosis each requires an independent model), which will cause storage redundancy and memory usage in the end, and will introduce too many dependent libraries, which is also a waste of resources.
[0003] In summary, in client scenarios, the existing methods of solving AI needs waste resources. Summary of the invention
[0004] The embodiment of the present invention provides an implementation scheme different from the related art to solve the technical problem in the related art that the existing method of solving AI needs in the client scenario wastes resources.
[0005] In a first aspect, the present invention provides a task processing method, comprising: In response to detecting a request to perform a first task, determining first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is performed by an artificial intelligence model deployed locally; Acquire first model parameter information corresponding to the first task category, where the first model parameter information is used to determine first model parameters of each computing unit in the target base model, and then generate a first model for performing the first task; Applying the first model parameter information to the target base model to obtain the first model; The first data to be processed is processed based on the first model to obtain an execution result of the first task.
[0006] In a second aspect, the present invention provides a task processing device, comprising: a determining unit, configured to determine, in response to detecting a request to perform a first task, first task information of the first task, the first task information comprising first data to be processed and a first task category of the first task, wherein the first task is performed by an artificial intelligence model deployed locally; an acquisition unit, configured to acquire first model parameter information corresponding to the first task category, wherein the first model parameter information is used to determine first model parameters of each computing unit in the target base model, and then generate a first model for performing the first task; an application unit, configured to apply the first model parameter information to the target base model to obtain the first model; A processing unit is used to process the first data to be processed based on the first model to obtain an execution result of the first task.
[0007] In a third aspect, the present invention provides an electronic device, comprising: Processor; and A memory, configured to store executable instructions of the processor; The processor is configured to execute the first aspect, or any method in any possible implementation manner of the first aspect, by executing the executable instructions.
[0008] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the first aspect, or any method in any possible implementation manner of the first aspect.
[0009] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the first aspect, or any method in any possible implementation manner of the first aspect.
[0010] The present invention provides a method for determining, in response to detecting a request to execute a first task, first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is executed by an artificial intelligence model deployed locally; obtaining first model parameter information corresponding to the first task category, the first model parameter information being used to determine first model parameters of each computing unit in a target base model, and then generating a first model for executing the first task; applying the first model parameter information to the target base model to obtain the first model; processing the first data to be processed based on the first model to obtain a solution for the execution result of the first task, wherein when executing the first task, the first model parameter information corresponding to the first task category of the first task can be applied to the target base model, and then determining the first model for executing the first task, the process of determining the first model is simple, saving the model layout cost, and saving storage and computing resources when solving AI needs based on local models. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 A schematic diagram of the structure of a system provided by an embodiment of the present invention; Figure 2a A flowchart of a task processing method provided by an embodiment of the present invention; Figure 2b A schematic diagram of a home page of a client provided by an embodiment of the present invention; Figure 2c According to an embodiment of the present invention, for each model parameter value in the first model parameter information, the model parameter value is used as an operation coefficient of a calculation unit corresponding to the model parameter value in the target base model to obtain a scene schematic diagram of the first model; Figure 2d A schematic diagram of a process of downloading a model parameter file and a configuration file corresponding to a first task category provided by an embodiment of the present invention; Figure 2e A flowchart of a task processing method provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a task processing device provided by an embodiment of the present invention; Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] Embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present invention, but should not be understood as limiting the present invention.
[0013] The terms "first" and "second" and the like in the specification, claims and drawings of the embodiments of the present invention are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0014] First, some terms in the embodiments of the present invention are explained below to facilitate understanding by those skilled in the art.
[0015] SDK, short for Software Development Kit, is a collection of files that includes development tools for creating applications for specific software packages, software frameworks, hardware platforms, operating systems, etc.
[0016] The base model is a concept widely used in the fields of deep learning and artificial intelligence. A base model is a basic model used to support other more complex models or to initialize model parameters. In deep learning, base models are usually pre-trained models, such as BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), etc. They are trained on large-scale corpora and have strong representation capabilities and generalization performance. These base models can be fine-tuned on specific tasks to adapt to different application scenarios. BERT is a pre-trained language representation model. The full name of BERT is "Bidirectional Encoder Representations from Transformers". It is based on the Transformer architecture. It learns the deep features of language by pre-training on large-scale unsupervised corpora, and then used in various supervised downstream natural language processing tasks. The Transformer architecture is a neural network architecture based on the attention mechanism, which was originally proposed for sequence-to-sequence learning in natural language processing tasks.
[0017] The inventors have discovered through research that currently, many services in some clients have or plan to use the on-end offline AI (Artificial Intelligence) reasoning capabilities to solve specific engineering problems on the end.
[0018] At present, some clients mainly use two methods to solve AI needs. One is to use online services (such as enterprise Q&A, meeting minutes, screenshot OCR) to solve AI needs, and the other is to use artificial algorithms, such as some AI conclusions cached locally to solve AI needs. When using these two solutions to deploy AI capabilities on the client, there are the following problems, which do not meet the growing requirements for on-end offline AI reasoning capabilities.
[0019] Problems with using online service reasoning: Compliance issues: High-level confidential user data cannot be sent directly to AI services for reasoning, which poses a greater compliance risk; Latency issue: Using online services to reason about specific problems requires complete interaction with the server. Large models take a long time to reason, and the client cannot obtain results in real time. Cost issue: General large models can accomplish a lot of work, but many daily business-specific problems do not require such large-scale, high-cost reasoning models all the time, which will lead to a waste of resources.
[0020] Problems with using local solution reasoning: Technical threshold issues: Some clients directly put model files into the end, which can solve specific needs but is not universal. In addition, the business side also needs to select, deploy, and adapt the AI model framework by themselves, which requires the business side to have high AI engineering capabilities.
[0021] Resource waste problem: In the client, multiple business scenarios independently deploy multiple models (such as phishing detection / network diagnosis each requires an independent model), which will cause storage redundancy and memory usage on the end, and will introduce too many dependent libraries, which will waste resources and easily cause security issues.
[0022] Although other methods in the industry have preset fixed general functional models and can solve certain problems, if the business needs to add new non-general capabilities, it is necessary to learn relevant knowledge and develop, train, and deploy complete models by yourself. The essential defect is that it fails to combine model structure reuse with business logic.
[0023] In certain scenarios, if the user's sensitive data cannot be checked in the cloud, it can only be inferred and calculated locally due to data compliance requirements. When such problems cannot be solved by manual algorithms, AI needs to be used, but there is currently no such unified solution in the company. This invention provides a universal solution to solve this demand.
[0024] Traditional cross-platform solutions require business parties to have professional AI deployment capabilities. Each business must develop its own AI and deploy it to other businesses. Businesses need to select models from scratch, which has high R&D costs and introduces complex code dependencies, increasing software maintenance costs.
[0025] The present invention relates to the field of artificial intelligence engineering technology, and is a system solution for rapidly deploying multi-service AI capabilities on mobile / embedded devices through a unified base model architecture combined with a dynamic parameter delivery mechanism. The present invention is implemented by a reconfigurable base model solidified on the device side and business model parameters dynamically delivered from the cloud, and the scheduling system connecting the two in the client completes the function.
[0026] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0027] The present invention can be applied to, for example, office software products, which have multiple functional modules, including but not limited to one or more of instant messaging (IM), video conferencing, shared documents, spreadsheets, mailboxes, calendars, and the like.
[0028] Figure 1 A structural diagram of a system provided for an exemplary embodiment of the present invention, the system includes a terminal 10 and a server 20, the terminal 10 may refer to a client device, the terminal 10 may be installed with multiple clients, in the present invention, the client refers to a client application, each client may have at least one built-in functional module, the at least one functional module may include at least one of the following: document module, audio and video conferencing module, calendar module, address book module, shared document module or instant messaging module. The user can use each functional module in at least one client, and can trigger the corresponding AI model processing task in the functional module. Specifically, the user can trigger one or more AI model processing tasks in the functional module, and the tasks processed by multiple AI intelligent models can be different. The tasks under different functional modules can be different. For example, when the functional module is a document module, the user can trigger the task of translating the document content or the task of polishing the document content in the functional module. When the functional module is an instant messaging module, the user can trigger the tasks of intelligent reply and intelligent chat in the functional module.
[0029] Furthermore, the terminal 10 is used for: In response to detecting a request to perform a first task, determining first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is performed by an artificial intelligence model deployed locally; Acquire first model parameter information corresponding to the first task category, where the first model parameter information is used to determine first model parameters of each computing unit in the target base model, and then generate a first model for performing the first task; Applying the first model parameter information to the target base model to obtain the first model; The first data to be processed is processed based on the first model to obtain an execution result of the first task.
[0030] The server 20 is used to send the latest first model parameter information corresponding to the first task category to the terminal 10 .
[0031] The execution principles and interaction processes of the various components in the system embodiment, such as the terminal 10 and the various units in the server 20, can be found in the description of the following method embodiments.
[0032] Figure 2aA flowchart of a task processing method provided by an exemplary embodiment of the present invention is provided. The method can be applied to a client device, such as the aforementioned terminal 10. The method at least includes the following steps: S201, in response to detecting a request to execute a first task, determining first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is executed by an artificial intelligence model deployed locally; The artificial intelligence model deployed locally may specifically refer to a local base model. The first task category of the first task refers to the purpose of processing the first data to be processed, such as translation, polishing, summarization, beautification, text and graphics, etc.
[0033] In some optional embodiments of the present invention, the first task is a task triggered by a user through any functional module in a client, the client is installed locally, the functional modules are classified based on functions, and at least one functional module in the client includes at least one of the following: a document module, an audio and video conferencing module, a calendar module, an address book module, a shared document module, or an instant messaging module.
[0034] Users can use various functional modules in the client and trigger corresponding AI model processing tasks in the functional modules.
[0035] In some optional embodiments of the present invention, the home page of the client may display entry buttons for each functional module, for example, Figure 2b As shown, Figure 2b A schematic diagram of a client homepage is provided for an embodiment of the present invention, wherein function module 1, function module 2, and function module 3 are entry buttons of three function modules. After a user clicks the entry button of a function module, a function usage page under the function module is displayed to the user.
[0036] In some optional embodiments of the present invention, the method further includes: upon detecting a request for executing the first task initiated by any functional module through a communication interface of the client, determining that the request for executing the first task is detected.
[0037] Optionally, after determining the first task information of the first task, the aforementioned communication interface can call the AI SDK interface to create the first task.
[0038] S202, obtaining first model parameter information corresponding to the first task category, where the first model parameter information is used to determine first model parameters of each computing unit in a target base model, and then generate a first model for performing the first task; The aforementioned computing units are specifically each node in the target base model, that is, each neuron.
[0039] The user can trigger one or more AI model processing tasks in the function module, and multiple AI intelligent model processing tasks can be different. The tasks under different function modules can be different. For example, when the function module is a document module, the user can trigger the task of translating the document content or polishing the document content in the function module. When the function module is an instant messaging module, the user can trigger the tasks of intelligent reply and intelligent chat in the function module.
[0040] When the first task category is different, the functions to be implemented by the model are different. The first task category corresponds one-to-one to the first model parameter information. Different first task categories can share the same base model to generate different task processing models. The aforementioned first model and the following second model are task processing models.
[0041] In some optional embodiments of the present invention, before obtaining the first model parameter information corresponding to the first task category, the method also includes: detecting whether an existing model parameter file corresponding to the first task category is stored locally, and if so, verifying the legitimacy of the existing model parameter file; if the existing model parameter file passes the legitimacy verification, using the model parameter information in the existing model parameter file as the first model parameter information corresponding to the first task category.
[0042] Optionally, the model parameter file only includes multiple model parameter values and does not include other content.
[0043] Optionally, when the version of the existing model parameter file is the latest version of the model parameter file, and the hash value of the existing model parameter file is a preset hash value, it is determined that the existing model parameter file passes the legality verification. Optionally, the latest version of the model parameter file and the preset hash value can be stored locally or obtained from a server. By verifying the legality of the existing model parameter file, the accuracy of the determined model-processed data can be further improved.
[0044] If the existing model parameter file passes the legality verification, the local AI SDK loads the existing model parameter file into the memory, and uses the model parameter information in the existing model parameter file as the first model parameter information corresponding to the first task category.
[0045] Specifically, the verification of the legitimacy of existing model parameter files can be achieved through the local AI SDK.
[0046] Specifically, the first model is a target base model after each calculation unit is assigned a calculation coefficient, wherein the calculation coefficient is a model parameter value. Optionally, the aforementioned first task and the following second task may belong to the same functional module or to different functional modules.
[0047] Optionally, the aforementioned target base model is the only base model deployed locally.
[0048] Optionally, the aforementioned target base model is one of the base models among multiple base models deployed locally.
[0049] Optionally, the aforementioned target base model may be included in the AI SDK. The aforementioned target base model may be written in any language, such as Rust, C++, etc., and built into the client device. Different base models have different structures.
[0050] The target base model does not include any model parameter values, the model parameter values are empty, the target base model only has basic network structure and calculation relationship, and the first model is a network structure with model parameter values.
[0051] In some optional embodiments of the present invention, multiple base models are deployed locally, the first task information also includes a target model identifier of the target base model applicable to the first task category, and the method also includes: determining the target base model corresponding to the target model identifier from the multiple base models based on the target model identifier.
[0052] Different base models have different model identifiers, and the same base model can be used to generate different task processing models.
[0053] S203, applying the first model parameter information to the target base model to obtain the first model; Optionally, the aforementioned S203 is also implemented by the aforementioned AI SDK.
[0054] In some optional embodiments of the present invention, the first model parameter information includes multiple model parameter values, and the first model parameter information is applied to the target base model to obtain the first model, including: for each model parameter value in the first model parameter information, applying the model parameter value to a calculation unit corresponding to the model parameter value in the target base model to obtain the first model, wherein the model parameter value corresponds one-to-one to the calculation unit.
[0055] In some optional embodiments of the present invention, for each model parameter value in the first model parameter information, the model parameter value is applied to the calculation unit corresponding to the model parameter value in the target base model to obtain the first model, including: for each model parameter value in the first model parameter information, the model parameter value is used as the operation coefficient of the calculation unit corresponding to the model parameter value in the target base model to obtain the first model.
[0056] Optionally, the first model can be regarded as a calculation function, assuming that the target base model is: f(x) = ?? x² + ?? x + ?? Optionally, x is a variable representing an input to the function.
[0057] When the first model parameter information is: -2, 8, 1, the first model is: f(x) = -2x² + 8x + 1.
[0058] When the first model parameter information is: 0, 1, 1, the first model is: f(x) = 0x² + 1x + 1 = x +1.
[0059] When the first model parameter information is: 0, 0, 4, the first model is: f(x) = 0x² + 0x + 4 = 4.
[0060] The aforementioned first model can be regarded as an AI function that can be called repeatedly. When the AI function is formed, the AI SDK passes the first data to be processed to the AI function, and the AI function locally executes the data processing process and returns the processing result to the functional module through the client. Optionally, the processing result can be displayed on the corresponding page.
[0061] Further, see Figure 2c As shown, Figure 2c For each model parameter value in the first model parameter information, the model parameter value is used as an operation coefficient of a calculation unit corresponding to the model parameter value in the target base model to obtain a scene schematic diagram of the first model.
[0062] Figure 2c The first model parameter information includes multiple model parameter values: 0.2, 0.3, 1.6, 0.6, -2.7, 0.8, 0.5, 0.2, ... 0.1, 0.3, 0.4, 0.5, 0.8, 0.2, 0.8. The target base model includes multiple calculation units: OP1, OP2, OP3, OP4, ... OPw, OPx, OPy, OPz, OPr, OPs, OPm, OPn, OPa, OPb, OP_final. Each model parameter value corresponds to a calculation unit. For each model parameter value in the first model parameter information, the model parameter value is applied to the calculation unit corresponding to the model parameter value in the target base model. The result of the first model can be seen in Figure 2c The first model is shown.
[0063] Among them, the "?" in the present invention is used to refer to space.
[0064] S204: Process the first data to be processed based on the first model to obtain an execution result of the first task.
[0065] Optionally, the input format of the first data to be processed is determined by the network structure of the first model, and the first data to be processed may be converted into a format corresponding to the first model before being processed based on the first model.
[0066] In some optional embodiments of the present invention, the method further includes the following steps S301-S304: S301, in response to detecting a request to execute a second task, determining second task information of the second task, where the second task information includes second data to be processed and a second task category of the second task, wherein the second task is also executed by a locally deployed artificial intelligence model; In some optional embodiments of the present invention, the second task is a task triggered by a user through any functional module in the client.
[0067] In some optional embodiments of the present invention, the method further includes: upon detecting a request for executing the second task initiated by any functional module through the communication interface of the client, determining that the request for executing the second task is detected.
[0068] S302, obtaining second model parameter information corresponding to the second task category, where the second model parameter information is used to determine second model parameters of each computing unit in the target base model, and then generate a second model for performing the second task; When the second task category is different, the functions to be implemented by the model are different. The second task category corresponds one-to-one to the second model parameter information. Different second task categories can share the same base model to generate different task processing models. The aforementioned first model and second model are task processing models.
[0069] S303, applying the second model parameter information to the target base model to obtain the second model; The first model and the second model are generated based on the same target base model.
[0070] Optionally, the principle of applying the second model parameter information to the target base model to obtain the second model is the same as the principle of applying the first model parameter information to the target base model to obtain the first model. The same is not repeated here.
[0071] S304: Process the second data to be processed based on the second model to obtain an execution result of the second task.
[0072] Through the above scheme, the target base model can also be used to generate a second model different from the first model to perform a second task. In the present invention, when performing tasks based on the model, the model determination process is simple, and different task execution models can share the internal structure, saving the model layout cost, and saving storage and computing resources when solving AI needs based on local models. In some optional embodiments of the present invention, before determining the first task information of the first task in response to detecting a request to perform the first task, the method also includes the following S001-S003: S001. Sending a configuration acquisition request for the first task category to a server; Optionally, the aforementioned S001 can be triggered after the client is started, and after detecting that the client is started, multiple task categories corresponding to the client can be obtained, and a configuration acquisition request for each task category can be sent to the server. Among them, the multiple task categories corresponding to the client refer to all task categories of tasks triggered by various functional modules in the client.
[0073] Optionally, the plurality of task categories include the aforementioned first task category and second task category.
[0074] Optionally, the aforementioned S001 may also be executed periodically, which is not limited by the present invention.
[0075] S002, receiving configuration information corresponding to the first task category issued by the server based on the configuration acquisition request, the configuration information including: a download address of a model parameter file corresponding to the first task category, a version number of the model parameter file corresponding to the first task category, and a hash value of the model parameter file corresponding to the first task category, wherein the first model parameter information corresponding to the first task category is included in the model parameter file corresponding to the first task category; The configuration information corresponding to the first task category issued by the server based on the configuration acquisition request is the latest configuration information corresponding to the first task category.
[0076] In some optional embodiments of the present invention, the configuration information may also include a base model identifier of the base model used by the first task category. Accordingly, the above method also includes: verifying whether the base model identifier is included in the preset model identifier set of the preset base model corresponding to the first task category. If so, determine to execute the following step S003, wherein the preset model identifier set may be stored locally.
[0077] Optionally, the base model used by the first task category includes the aforementioned target base model.
[0078] By verifying the base model identification, errors in matching the first task category with the configuration information can be avoided, thereby improving the accuracy of determining the configuration information.
[0079] S003. Determine whether the hash value of the model parameter file corresponding to the first task category is the same as the hash value of the existing model parameter file corresponding to the first task category locally. If so, determine whether the version number of the model parameter file corresponding to the first task category is newer than the version number of the existing model parameter file. If so, download the configuration file corresponding to the first task category and the model parameter file corresponding to the first task category from the server based on the download address, wherein the configuration file includes the configuration information.
[0080] Among them, the configuration file and model parameter file that have been downloaded and stored locally are the existing configuration file and the existing model parameter file.
[0081] Optionally, if the hash value of the model parameter file corresponding to the first task category is different from the hash value of the existing model parameter file corresponding to the first task category locally, it will not be processed, or an error prompt will be sent to the target device to make the target device re-determine the configuration information corresponding to the first task category.
[0082] In some optional embodiments of the present invention, the configuration information further includes: a local file operation instruction, the local file operation instruction includes downloading or disabling, and the local file operation instruction is used to indicate a processing method for the existing configuration file and the existing model parameter file corresponding to the first task category locally; If the local file operation indication is downloading, then the execution of determining whether the hash value of the model parameter file corresponding to the first task category is the same as the hash value of the existing model parameter file corresponding to the first task category locally is triggered; if the local file operation indication is disabled, then the existing configuration file and the existing model parameter file are deleted.
[0083] Optionally, when an abnormality is detected on the client or in the execution of any task, the relevant personnel can set the local file operation indication to be disabled, which can avoid the failure of task execution and improve the reliability of task execution.
[0084] The multiple base models in the present invention can be stored locally and can be updated according to the model update instructions of relevant personnel.
[0085] Optionally, the download process of the model parameter file and configuration file corresponding to the first task category can also be found in Figure 2d As shown, Figure 2dA schematic diagram of a process of downloading a model parameter file and a configuration file corresponding to a first task category provided in an embodiment of the present invention. Specifically, the method further includes the following steps S1-S7: S1, requesting the server to send the latest configuration information of the first task category; S2. Receive the latest configuration information of the first task category sent by the server; S3, check whether the model parameter information corresponding to the first task category is updated, if so, execute the following step S4; if the model parameter information corresponding to the first task category is not updated, end.
[0086] S4, reading the content of the local file operation indication in the latest configuration information, determining whether the content of the local file operation indication is download or disable, if the content of the local file operation indication is download, executing step S5; if the content of the local file operation indication is disable, executing the following step S6; S5, downloading the model parameter file and configuration file corresponding to the first task category from the server; S6. Delete the existing local model parameter file and the existing configuration file corresponding to the first task category.
[0087] S7. End.
[0088] In order to further illustrate the solution of the present invention, Figure 2e A flowchart of a task processing method is also provided. The method comprises the following steps S11-S15: S11. The service initiates a request to execute the first task based on the local base model through the unified communication interface of the client; wherein the service refers to the functional module.
[0089] S12. The AI SDK loads a model parameter file corresponding to the first task category of the first task; S13. The AI SDK applies the first model parameter information in the model parameter file to the target base model corresponding to the first task category of the first task to obtain a first model. S14, executing the first task based on the first model to obtain an execution result; S15. AI SDK transmits the execution result to the client's unified communication interface and returns it to the business.
[0090] Optionally, the communication interface may refer to JSBridge API (Application Programming Interface). JSBridge API plays a vital role in hybrid development, enabling two-way communication between the Web end (H5) and the native end.
[0091] The solution of the present invention provides the client with a unified, easy-to-use, offline inference AI system within the client. While ensuring functional scalability, it only uses the processor to infer based on the AI model, controls the low memory usage of the client device during AI inference, and supports offline inference without a network environment. In addition, the physical separation mechanism of the model structure and business parameters enables the same base model to perform multiple tasks through different parameter configurations, avoiding the multi-model redundancy of traditional solutions, and can also effectively protect the model within the terminal from being stolen; AI model selection only requires parameter selection, which greatly improves efficiency compared to traditional AI model selection.
[0092] The present invention completes the function use through a reconfigurable base model solidified in a device end, model parameters dynamically sent down by a server, and a scheduling system connecting the base model and the model parameters in a client end.
[0093] The present invention provides a method for determining, in response to detecting a request to execute a first task, first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is executed by an artificial intelligence model deployed locally; obtaining first model parameter information corresponding to the first task category, the first model parameter information being used to determine first model parameters of each computing unit in a target base model, and then generating a first model for executing the first task; applying the first model parameter information to the target base model to obtain the first model; processing the first data to be processed based on the first model to obtain a solution for the execution result of the first task, wherein when executing the first task, the first model parameter information corresponding to the first task category of the first task can be applied to the target base model, and then determining the first model for executing the first task, the process of determining the first model is simple, saving the model layout cost, and saving storage and computing resources when solving AI needs based on local models.
[0094] Figure 3 A schematic diagram of a data processing device provided by an exemplary embodiment of the present invention, wherein the device comprises: a determining unit 31, configured to determine, in response to detecting a request to perform a first task, first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is performed by an artificial intelligence model deployed locally; An acquisition unit 32 is used to acquire first model parameter information corresponding to the first task category, where the first model parameter information is used to determine first model parameters of each computing unit in the target base model, and then generate a first model for performing the first task; An application unit 33, configured to apply the first model parameter information to the target base model to obtain the first model; The processing unit 34 is used to process the first data to be processed based on the first model to obtain an execution result of the first task.
[0095] In some optional embodiments of the present invention, The determining unit 31 is further configured to determine, in response to detecting a request to perform a second task, second task information of the second task, the second task information including second data to be processed and a second task category of the second task, wherein the second task is also performed by the artificial intelligence model deployed locally; The acquisition unit 32 is further used to acquire second model parameter information corresponding to the second task category, where the second model parameter information is used to determine second model parameters of each computing unit in the target base model, and then generate a second model for performing the second task; The application unit 33 is further configured to apply the second model parameter information to the target base model to obtain the second model; The processing unit 34 is further configured to process the second data to be processed based on the second model to obtain an execution result of the second task.
[0096] In some optional embodiments of the present invention, multiple base models are deployed locally, the first task information also includes a target model identifier of the target base model applicable to the first task category, and the device is further used to: determine the target base model corresponding to the target model identifier from the multiple base models based on the target model identifier.
[0097] In some optional embodiments of the present invention, the first model parameter information includes multiple model parameter values. When the device is used to apply the first model parameter information to the target base model to obtain the first model, it is specifically used to: for each model parameter value in the first model parameter information, apply the model parameter value to the calculation unit corresponding to the model parameter value in the target base model to obtain the first model, wherein the model parameter value corresponds to the calculation unit one-to-one.
[0098] In some optional embodiments of the present invention, when the aforementioned apparatus is used to apply, for each model parameter value in the first model parameter information, the model parameter value to a calculation unit corresponding to the model parameter value in the target base model to obtain the first model, it is specifically used to: For each model parameter value in the first model parameter information, the model parameter value is used as an operation coefficient of a calculation unit corresponding to the model parameter value in the target base model to obtain the first model.
[0099] In some optional embodiments of the present invention, the aforementioned device is further used, before being used to obtain the first model parameter information corresponding to the first task category, to: detect whether an existing model parameter file corresponding to the first task category is stored locally; if so, verify the legitimacy of the existing model parameter file; if the existing model parameter file passes the legitimacy verification, use the model parameter information in the existing model parameter file as the first model parameter information corresponding to the first task category.
[0100] In some optional embodiments of the present invention, before determining the first task information of the first task in response to detecting a request to perform the first task, the aforementioned apparatus is further configured to: Sending a configuration acquisition request for the first task category to a server; Receive configuration information corresponding to the first task category issued by the server based on the configuration acquisition request, the configuration information including: a download address of a model parameter file corresponding to the first task category, a version number of the model parameter file corresponding to the first task category, and a hash value of the model parameter file corresponding to the first task category, wherein the first model parameter information corresponding to the first task category is included in the model parameter file corresponding to the first task category; Determine whether the hash value of the model parameter file corresponding to the first task category is the same as the hash value of the existing model parameter file corresponding to the first task category locally; if so, determine whether the version number of the model parameter file corresponding to the first task category is newer than the version number of the existing model parameter file; if so, download the configuration file corresponding to the first task category and the model parameter file corresponding to the first task category from the server based on the download address, wherein the configuration file includes the configuration information.
[0101] In some optional embodiments of the present invention, the configuration information further includes: a local file operation instruction, the local file operation instruction includes downloading or disabling, and the local file operation instruction is used to indicate a processing method for the existing configuration file and the existing model parameter file corresponding to the first task category locally; The aforementioned device is also used for: if the local file operation indication is downloading, triggering the execution of determining whether the hash value of the model parameter file corresponding to the first task category is the same as the hash value of the existing model parameter file corresponding to the first task category locally; if the local file operation indication is disabled, deleting the existing configuration file and the existing model parameter file.
[0102] In some optional embodiments of the present invention, the first task is a task triggered by a user through any functional module in a client, the client is installed locally, the functional modules are classified based on functions, and at least one functional module in the client includes at least one of the following: a document module, an audio and video conferencing module, a calendar module, an address book module, a shared document module, or an instant messaging module.
[0103] In some optional embodiments of the present invention, the aforementioned apparatus is further used to: upon detecting a request for executing the first task initiated by any functional module through the communication interface of the client, determine that the request for executing the first task is detected.
[0104] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, the device may perform the above method embodiment, and the above and other operations and / or functions of each module in the device are the corresponding processes in each method in the above method embodiment, respectively, and no further description is given here for the sake of brevity.
[0105] The above describes the device of the embodiment of the present invention from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present invention can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present invention can be directly embodied as a hardware decoding processor to execute, or a combination of hardware and software modules in the decoding processor to execute. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.
[0106] Figure 4 is a schematic block diagram of an electronic device provided by an embodiment of the present invention, and the electronic device may include: The memory 401 and the processor 402, the memory 401 is used to store the computer program and transmit the program code to the processor 402. In other words, the processor 402 can call and run the computer program from the memory 401 to implement the method in the embodiment of the present invention.
[0107] For example, the processor 402 may be configured to execute the above method embodiments according to instructions in the computer program.
[0108] In some embodiments of the present invention, the processor 402 may include but is not limited to: General-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0109] In some embodiments of the present invention, the memory 401 includes but is not limited to: Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct memory bus random access memory (Direct Rambus RAM, DR RAM).
[0110] In some embodiments of the present invention, the computer program may be divided into one or more modules, which are stored in the memory 401 and executed by the processor 402 to complete the method provided by the present invention. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0111] like Figure 4As shown, the electronic device may further include: a transceiver 403 , which may be connected to the processor 402 or the memory 401 .
[0112] The processor 402 may control the transceiver 403 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 403 may include a transmitter and a receiver. The transceiver 403 may further include an antenna, and the number of antennas may be one or more.
[0113] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0114] The present invention also provides a computer storage medium on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the present invention also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.
[0115] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0116] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0117] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the module is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0118] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, the functional modules in various embodiments of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0119] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A task processing method, characterized in that: include: In response to detecting a request to perform a first task, determining first task information of the first task, the first task information including first data to be processed and a first task category of the first task, wherein the first task is performed by an artificial intelligence model deployed locally; Acquire first model parameter information corresponding to the first task category, where the first model parameter information is used to determine first model parameters of each computing unit in the target base model, and then generate a first model for performing the first task; Applying the first model parameter information to the target base model to obtain the first model; The first data to be processed is processed based on the first model to obtain an execution result of the first task.
2. The task processing method according to claim 1, characterized in that: The method further comprises: In response to detecting a request to perform a second task, determining second task information of the second task, the second task information including second data to be processed and a second task category of the second task, wherein the second task is also performed by the locally deployed artificial intelligence model; Acquire second model parameter information corresponding to the second task category, where the second model parameter information is used to determine second model parameters of each computing unit in the target base model, and then generate a second model for performing the second task; Applying the second model parameter information to the target base model to obtain the second model; The second data to be processed is processed based on the second model to obtain an execution result of the second task.
3. The task processing method according to claim 1, characterized in that: There are multiple base models deployed locally, the first task information also includes a target model identifier of the target base model applicable to the first task category, and the method also includes: based on the target model identifier, determining the target base model corresponding to the target model identifier from the multiple base models.
4. The task processing method according to claim 1, characterized in that: The first model parameter information includes a plurality of model parameter values, and applying the first model parameter information to the target base model to obtain the first model includes: For each model parameter value in the first model parameter information, the model parameter value is applied to a calculation unit corresponding to the model parameter value in the target base model to obtain the first model, wherein the model parameter value corresponds to the calculation unit one-to-one.
5. The task processing method according to claim 4, characterized in that: The step of applying, for each model parameter value in the first model parameter information, the model parameter value to a calculation unit corresponding to the model parameter value in the target base model to obtain the first model includes: For each model parameter value in the first model parameter information, the model parameter value is used as an operation coefficient of a calculation unit corresponding to the model parameter value in the target base model to obtain the first model.
6. The task processing method according to claim 1, characterized in that: Before obtaining the first model parameter information corresponding to the first task category, the method also includes: detecting whether an existing model parameter file corresponding to the first task category is stored locally, and if so, verifying the legitimacy of the existing model parameter file; if the existing model parameter file passes the legitimacy verification, using the model parameter information in the existing model parameter file as the first model parameter information corresponding to the first task category.
7. The task processing method according to claim 1, characterized in that: Before determining first task information of the first task in response to detecting a request to perform the first task, the method further includes: Sending a configuration acquisition request for the first task category to a server; Receive configuration information corresponding to the first task category issued by the server based on the configuration acquisition request, the configuration information including: a download address of a model parameter file corresponding to the first task category, a version number of the model parameter file corresponding to the first task category, and a hash value of the model parameter file corresponding to the first task category, wherein the first model parameter information corresponding to the first task category is included in the model parameter file corresponding to the first task category; Determine whether the hash value of the model parameter file corresponding to the first task category is the same as the hash value of the existing model parameter file corresponding to the first task category locally; if so, determine whether the version number of the model parameter file corresponding to the first task category is newer than the version number of the existing model parameter file; if so, download the configuration file corresponding to the first task category and the model parameter file corresponding to the first task category from the server based on the download address, wherein the configuration file includes the configuration information.
8. The task processing method according to claim 7, characterized in that: The configuration information further includes: a local file operation instruction, the local file operation instruction includes downloading or disabling, and the local file operation instruction is used to indicate a processing method for the existing configuration file and the existing model parameter file corresponding to the first task category locally; If the local file operation indication is downloading, then the execution of determining whether the hash value of the model parameter file corresponding to the first task category is the same as the hash value of the existing model parameter file corresponding to the first task category locally is triggered; if the local file operation indication is disabled, then the existing configuration file and the existing model parameter file are deleted.
9. The task processing method according to claim 1, characterized in that: The first task is a task triggered by a user through any functional module in the client, the client is installed locally, the functional modules are classified based on functions, and at least one functional module in the client includes at least one of the following: a document module, an audio and video conferencing module, a calendar module, an address book module, a shared document module, or an instant messaging module.
10. The task processing method according to claim 9, characterized in that: The method further comprises: When detecting a request for executing the first task initiated by any functional module through the communication interface of the client, it is determined that the request for executing the first task is detected.
11. A task processing device, characterized in that: include: a determining unit, configured to determine, in response to detecting a request to perform a first task, first task information of the first task, the first task information comprising first data to be processed and a first task category of the first task, wherein the first task is performed by an artificial intelligence model deployed locally; an acquisition unit, configured to acquire first model parameter information corresponding to the first task category, wherein the first model parameter information is used to determine first model parameters of each computing unit in the target base model, and then generate a first model for performing the first task; an application unit, configured to apply the first model parameter information to the target base model to obtain the first model; A processing unit is used to process the first data to be processed based on the first model to obtain an execution result of the first task.
12. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the task processing method according to any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the task processing method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Task processing method and automatic question answering method
CN117193964A
Task scheduling method and device, computer equipment, storage medium and program product
CN117742970A
Model file determination method and device, storage medium and electronic equipment
CN118381787A
Task processing method and task processing system
CN119336477A
Model calculation scheduling method and device, equipment, medium and product
CN119806814A