Artificial intelligence-assisted trial methods, systems, equipment, storage media and products

By designing an artificial intelligence-assisted trial system, using multiple service models of the scenario application layer, model application layer and model service layer, the resource waste caused by single task processing in the existing technology is solved, and the efficiency and accuracy of judicial trials are achieved.

CN119477610BActive Publication Date: 2025-08-22SHENZHEN DIBO ENTERPRISE RISK MANAGEMENT TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510055774.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-08-22
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In the field of justice, artificial intelligence models can only handle a single task, resulting in waste of resources and inefficiency, and are unable to effectively assist judges in complex legal documents.

Method used

An artificial intelligence-assisted trial system was designed, including the scenario application layer, the model application layer and the model service layer. By obtaining case information and business scenario information, multiple service models are used for inference processing, and matching and processing of different types of tasks are achieved.

Benefits of technology

It improves the efficiency and accuracy of judicial trials, can dynamically schedule according to task type and model instance status, adapt to various judicial needs, and provide efficient legal reasoning and Q&A support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477610B_ABST
    Figure CN119477610B_ABST
Patent Text Reader

Abstract

The present disclosure provides an artificial intelligence-assisted trial method, system, device, storage medium, and product. The system includes a scenario application layer, a model application layer, and a model service layer. The scenario application layer is used to obtain case information and business scenario information of a task. The model application layer is used to obtain task configuration parameters of the task based on the business scenario information. The model service layer is used to obtain model instance running status information and obtain case information and task configuration parameters of the task; obtain input text vectors of the case information and task configuration parameters of the task, and model capability vectors of different types of service models; obtain the matching degree between the task and different types of service models based on the input text vector and model capability vector; distribute the case information and task configuration parameters of the task to the corresponding model instance of the corresponding type of service model for processing based on the matching degree and the model instance running status information, and obtain the reasoning result of the task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-assisted trial system, an artificial intelligence-assisted trial method, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In recent years, the application of artificial intelligence and deep learning has become increasingly widespread across various verticals. In the judicial field, AI research primarily focuses on legal theory, ethical issues, and algorithmic risks. The deep integration of law and technology to create intelligent applications of logical reasoning and decision-making capabilities is a practical need for smart justice. Faced with massive amounts of complex legal documents, relevant technologies can only apply single models to single tasks, resulting in a waste of resources. Summary of the Invention

[0003] An embodiment of the present disclosure provides an artificial intelligence-assisted trial system, which includes: a scenario application layer, which is used to obtain case information and business scenario information of a task; a model application layer, which is used to obtain case information and business scenario information of a task from the scenario application layer, and obtain task configuration parameters of the task based on the business scenario information; a model service layer, in which different types of service models are configured, and the same type of service model runs at least one model instance, and the model service layer is used to obtain model instance running status information, and obtain case information and task configuration parameters of the task from the model application layer; obtain input text vectors of the case information of the task and the task configuration parameters, and obtain model capability vectors of different types of service models; obtain the matching degree between the task and different types of service models based on the input text vector and the model capability vector; distribute the case information and task configuration parameters of the task to the corresponding model instance of the corresponding type of service model for processing based on the matching degree and the model instance running status information, and obtain the reasoning result of the task for assisting in the trial of the corresponding case.

[0004] An embodiment of the present disclosure provides an artificial intelligence-assisted trial method, which includes: obtaining case information and business scenario information of a task; obtaining task configuration parameters of the task based on the business scenario information; obtaining model instance running status information; obtaining input text vectors of the case information of the task and the task configuration parameters, and obtaining model capability vectors of different types of service models; obtaining the degree of matching between the task and different types of service models based on the input text vector and the model capability vector; and distributing the case information of the task and the task configuration parameters to corresponding model instances of corresponding types of service models for processing based on the matching degree and the model instance running status information, to obtain an inference result of the task for use in assisting the trial of the corresponding case.

[0005] An embodiment of the present disclosure provides a computer device, including a processor, a memory, and an input / output interface; the processor is connected to the memory and the input / output interface, respectively, wherein the input / output interface is used to receive and output data, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device including the processor executes the method in any embodiment of the present disclosure.

[0006] An embodiment of the present disclosure provides a computer-readable storage medium storing a computer program. The computer program is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method in any embodiment of the present disclosure.

[0007] The present disclosure provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in any of the various optional embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 It is a schematic diagram of an artificial intelligence-assisted trial system provided by an embodiment of the present disclosure.

[0009] Figure 2 It is a schematic diagram of another artificial intelligence-assisted trial system provided by an embodiment of the present disclosure.

[0010] Figure 3 This is a schematic diagram of another artificial intelligence-assisted trial system provided by an embodiment of the present disclosure.

[0011] Figure 4 Schematic diagram of the output of a detection model provided by an embodiment of the present disclosure.

[0012] Figure 5 It is a schematic diagram of the output of an entity extraction model provided by an embodiment of the present disclosure.

[0013] Figure 6 Schematic diagram of the output of a generative model provided by an embodiment of the present disclosure.

[0014] Figure 7 This is a schematic diagram of another artificial intelligence-assisted trial system provided by an embodiment of the present disclosure.

[0015] Figure 8 This is a schematic diagram of another artificial intelligence-assisted trial system provided by an embodiment of the present disclosure.

[0016] Figure 9 It is a schematic diagram of the formatting of entity extraction results provided by an embodiment of the present disclosure.

[0017] Figure 10 This is a flowchart of an artificial intelligence-assisted trial method provided by an embodiment of the present disclosure.

[0018] Figure 11 This is a network architecture diagram for implementing an artificial intelligence-assisted trial system provided by an embodiment of the present disclosure.

[0019] Figure 12 It is a structural diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0021] First, some terms involved in the embodiments of the present disclosure are explained.

[0022] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0023] In the embodiments disclosed herein, machine learning technology in artificial intelligence is used to assist court staff, such as judges, in conducting judicial trials to improve trial efficiency and accuracy.

[0024] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI.

[0025] like Figure 1 As shown, the artificial intelligence assisted trial system 10 provided in the embodiment of the present disclosure includes a scenario application layer 110, a model application layer 120 and a model service layer 130.

[0026] The scenario application layer 110 is used to obtain task case information and business scenario information.

[0027] The scenario application layer in the embodiment of the present disclosure refers to an application layer used to receive original case materials and / or extract elements from the received original case materials. It can also be used to interact with the front end, and according to the task request sent from the front end, obtain the case information corresponding to the request from the original case materials and / or the extracted case element information, and can determine the corresponding business scenario information based on the request.

[0028] In the embodiments of the present disclosure, a task is triggered by a user in the AI-assisted review system to request the completion of a task specified by the user and completed using a service model. Therefore, it can also be called a model task. The case information of a task refers to any data and / or information related to the pending case corresponding to the task. Business scenario information refers to the type of task to be completed based on the request. Based on the business scenario information, it can be determined what type of service model to be called.

[0029] The model application layer 120 is used to obtain case information and business scenario information of the task from the scenario application layer 110, and obtain task configuration parameters of the task according to the business scenario information.

[0030] In some embodiments, the task configuration parameters include a model identifier corresponding to the business scenario information.

[0031] The model application layer in the disclosed embodiments refers to the application layer that interacts with the scenario application layer and the model service layer to determine how to schedule tasks received from the scenario application layer and send the scheduled tasks to the model service layer. The model application layer can pre-configure task configuration parameters corresponding to various business scenario information. When the model application layer receives a task request from the scenario application layer, it can retrieve the corresponding task configuration parameters from it.

[0032] The task configuration parameters in the embodiment of the present disclosure refer to configuration parameters related to the task, which are used to indicate the scheduling of the task, the service model required to be adopted by the task, and the relevant parameters of the task when performing reasoning using the service model.

[0033] In the embodiments of this disclosure, model identifiers are used to uniquely identify various types of service models in the model service layer. This model identifier can be the identity (ID) / index of a service model or the name of a service model, which is not limited in this disclosure. The following embodiments use model IDs as examples, but this disclosure is not limited to this.

[0034] The model service layer 130 is configured with different types of service models, and the same type of service model runs at least one model instance. The model service layer is used to obtain the model instance running status information, and obtain the case information of the task and the task configuration parameters from the model application layer; obtain the input text vector of the case information of the task and the task configuration parameters, and obtain the model capability vectors of different types of service models; obtain the matching degree between the task and different types of service models based on the input text vector and the model capability vector; distribute the case information of the task and the task configuration parameters to the corresponding model instance of the corresponding type of service model for processing based on the matching degree and the model instance running status information, and obtain the reasoning result of the task to assist in the trial of the corresponding case.

[0035] In the embodiment of the present disclosure, the model service layer refers to an application layer for providing a variety of different types of service models for performing reasoning services for various tasks.

[0036] The service model in the embodiments of this disclosure refers to an artificial intelligence model, machine learning model, or deep learning model that can be used to assist in judicial trials. In the following embodiments, the model service layer includes two types of service models: a deep learning model and a generative model. However, this disclosure is not limited to this. Other types of service models can be added to the model service layer based on the specific needs of judicial trials.

[0037] In the disclosed embodiment, the model service layer can open one or more model instances for each type of service model as needed, and each model instance can be used to process tasks of a corresponding type. For example, a model instance of a summary model can be used to process the task of generating a summary.

[0038] In the embodiment of the present disclosure, the model instance running status information refers to the load, available resources, whether there is a fault, and other related information of each model instance.

[0039] In some embodiments, when the task includes a question, the system is further used to: convert the question into a question vector; obtain a legal text vector of laws, regulations and / or cases; calculate the semantic similarity between the question vector and the legal text vector; use the laws, regulations and / or cases, and the question as graph nodes, obtain the edge weights between the graph nodes based on the semantic similarity between the graph nodes and the preset logical relationship of the trial ideas, use the question as the starting node, and find the reasoning chain through graph search; obtain the credibility of the reasoning chain; obtain the priority index of the candidate legal text based on the semantic similarity between the question vector and the legal text vector, and the credibility of the reasoning chain; determine the laws, regulations and / or cases corresponding to the question based on the priority index.

[0040] The disclosed embodiment can match legal text with question content through natural language processing and semantic modeling to achieve efficient legal reasoning and question answering. , and / or, Case Library , where n represents the number of laws and regulations, and m represents the number of cases in the case library. Both n and m are positive integers greater than or equal to 1, that is, the legal texts included in the legal text set can be laws and regulations and / or cases. Get the natural language description Q of the legal issue. According to the issue Q, find the relevant legal provisions or cases , and derive the legal reasoning path or reasoning chain P.

[0041] Each legal text is converted into a high-dimensional vector through an encoder, which is called a legal text vector. or , where f and c are any laws, regulations or cases in the legal text set; convert the legal issue into a problem vector, i.e. For example, the encoder is a large model specially trained using legal corpus.

[0042] In some embodiments, the semantic similarity between the legal text vector and the question vector can be calculated by the following formula :

[0043] (1)

[0044] Sort by semantic similarity and select the k entries with the highest semantic similarity as the preliminary selection results, where k is a positive integer greater than or equal to 1.

[0045] In the embodiment of the present disclosure, a legal reasoning path or reasoning chain P can also be constructed based on legal reasoning combined with logical relationships, which can be achieved through graph representation.

[0046] Specifically, legal provisions and / or cases are used as graph nodes. For example, the preliminary results of the k laws, regulations, and / or cases mentioned above can be used as graph nodes. Question Q is used as an input node, which is also a graph node. Next, the edge weights between different graph nodes are determined.

[0047] In an exemplary embodiment, the semantic similarity between different graph nodes can be calculated and used as the semantic similarity weight. For example, the semantic similarity weight can be calculated using the following formula: : .in, represents the semantic similarity weight between the i-th graph node and the j-th graph node; and are vectors representing the i-th graph node and the j-th graph node respectively.

[0048] In the embodiment of the present disclosure, logical relationships are preset based on the trial ideas provided by the judge, such as prior knowledge such as the trial idea document provided by the judge, and similarity judgments are made based on the trial ideas of the judge or other information provided to determine the weights of the logical relationships between different graph nodes.

[0049] For example, the edge weights between different graph nodes, i.e., the comprehensive weights, are determined based on the semantic similarity weights and the logical relationship weights: .in, represents the edge weight between the i-th graph node and the j-th graph node; Represents the logical relationship weight between the i-th graph node and the j-th graph node; and Respectively represent the importance of semantic similarity weight and logical relationship weight, and Both are greater than or equal to 0 and less than or equal to 1. Starting from Q, a graph search algorithm is used to find a chain of reasoning, P. Based on logical relationships (e.g., "causality" and "inclusion"), a multi-step reasoning path is constructed. This allows the corresponding laws, regulations, and / or classic cases for each chain of reasoning to be identified as candidate entries.

[0050] In an exemplary embodiment, the priority index of each candidate entry can be calculated according to the following formula :

[0051] (2)

[0052] In the above formula, , They are and The weights are all greater than or equal to 0 and less than or equal to 1. The credibility of the reasoning chain is calculated based on the reliability of the logical relationship in path P. Based on the priority index of each candidate entry, the candidate entry with the largest priority index is determined as the target entry, that is, the optimal legal clause and / or case is determined and returned as one of the response results to the question.

[0053] The artificial intelligence-assisted trial system provided by the embodiments of the present disclosure provides a general artificial intelligence-assisted system for various types of tasks for cases that may be encountered in judicial trial work. The system can configure corresponding task configuration parameters for various tasks, and can determine what type of model service to use to process the case information of the task based on the task configuration parameters and case information. It can also determine which model instance to use to process the task based on the model instance running status information. That is, no matter what type of task is provided, the service model suitable for processing the task can be determined based on the business scenario information, thereby improving the efficiency and accuracy of judicial trials.

[0054] In an exemplary embodiment, as Figure 2 As shown, the AI-assisted trial system 10 further includes a front end 140. The front end 140 is used to display an AI-assisted trial interface 141. The AI-assisted trial interface 141 includes a task trigger control (not shown). In response to a trigger operation on the corresponding task trigger control, the AI-assisted trial interface 141 sends a request for the task to the scenario application layer 110. The request includes a case identifier and request parameters. The scenario application layer 110 is used to determine a business scenario identifier based on the request parameters, and obtain corresponding case information based on the business scenario identifier and the case identifier, where the business scenario information includes the business scenario identifier.

[0055] In the disclosed embodiments, the AI-assisted trial interface can display case information and various task trigger controls. Different task trigger controls can trigger different tasks, such as generating a case summary or extracting the disputed issues. The disclosure does not limit the form, triggering method, or layout of these task trigger controls on the interface. For example, they can be triggered by voice, gesture, or by clicking a virtual button on the screen.

[0056] In the embodiment of the present disclosure, when a user triggers a specific task trigger control, the system can determine the corresponding case identifier based on which case is currently displayed on the interface. The case identifier can uniquely identify the corresponding case, and its specific form is not limited.

[0057] In the disclosed embodiment, each task trigger control is pre-assigned a corresponding request parameter, which is used to indicate how the user wishes to handle the case. In other embodiments, the user can also enter the request parameter through the AI-assisted trial interface 141 to indicate how the user wishes to handle the case.

[0058] In the embodiment of the present disclosure, the business scenario identifier (business scenario ID) is used to indicate the type of the task, or is a unique identifier of what kind of processing the user wants to perform on the case.

[0059] In the exemplary embodiment, continue to refer to Figure 2 The scenario application layer 110 includes a case filing module 111 , a middle platform module 112 , an elementization module 113 and an intelligent software development kit (SDK) 114 .

[0060] The case filing module 111 is used to receive case materials and perform structured processing on the case materials to obtain structured case data. In an exemplary embodiment, the case data includes one or more of litigation opinions, court trial transcripts, judgment documents, and evidence materials.

[0061] The elementization module 113 is used to obtain the case materials from the case filing module 111, extract elements from the case materials, and obtain case element information. In an exemplary embodiment, the case element information includes one or more of party information, evidence text, and case cause.

[0062] The middle platform module 112 is used to receive the request from the front end 140 and send the request to the smart SDK 114 .

[0063] Intelligent SDK 114 is configured to determine the business scenario identifier based on the request parameters, obtain corresponding case data from case filing module 111 and / or obtain corresponding case element information from elementization module 113 based on the business scenario identifier and the case identifier, and transmit the obtained case data and / or case element information to model application layer 120. The case information includes the case data and / or the case element information.

[0064] In the exemplary embodiment, continue to refer to Figure 2The model application layer 120 includes a preprocessor 121 , a prompt word manager 122 , an asynchronous model priority queue 123 , a real-time model priority queue 124 and a scheduler 125 .

[0065] The preprocessor 121 is configured to obtain the case information and business scenario information of the task from the scenario application layer 110 , and send the business scenario information to the prompt word manager 122 .

[0066] In exemplary embodiments, in the AI-assisted trial system architecture provided by the embodiments of the present disclosure, the input information obtained by the preprocessor from the scenario application layer is mainly structured case data or extracted element information after being processed by the scenario application layer. This may include one or more of the following:

[0067] (1) Case-related text information (i.e., case data): such as litigation opinions, court transcripts, judgment documents, evidence materials, etc. This text information is further cleaned, segmented, or parsed to adapt to the input format of subsequent models.

[0068] (2) Case element information: The scenario application layer extracts key information such as party information, evidence text, case cause, etc. through the elementization module.

[0069] (3) Request parameters: Determined based on the specific request input by the user, such as analysis of litigation content, generation of judgment documents, etc.

[0070] (4) Business scenario ID representing a specific scenario application.

[0071] The information obtained by the preprocessor from the scenario application layer is a structured case text data (i.e., case data), and carries the corresponding business scenario ID according to different scenario requirements, so that the preprocessor can clean the data, convert the format, enhance the data, and convert it into the parameters required for the prompt words for subsequent model inference steps.

[0072] The prompt word manager 122 is used to receive the business scenario information, determine the corresponding model identification, task priority and timeliness configuration based on the business scenario information, and send the determined model identification, task priority and timeliness configuration to the preprocessor 121, where the task configuration parameters include the model identification, task priority and timeliness configuration.

[0073] The preprocessor 121 is further configured to push the task to the asynchronous model priority queue 123 or the real-time model priority queue 124 according to the task priority and timeliness configuration.

[0074] The real-time priority queue and the asynchronous priority queue are two task management mechanisms used in the system to schedule and process different types of tasks. The main difference between these two queues lies in the timeliness and priority of task processing, and they essentially serve different business needs. The following is a detailed description of the two and their essential differences:

[0075] (1) Real-time model priority queue (also called real-time queue):

[0076] The real-time model priority queue is used to process tasks with high timeliness requirements (referred to as real-time tasks). These tasks require immediate responses, and the processing results (i.e., inference results) have a significant impact on user experience or business processes. Tasks in this queue have the following characteristics:

[0077] High timeliness: Tasks in the real-time queue require the system to be processed quickly and return results in a very short time. Task response time is a priority.

[0078] Higher priority: When allocating system resources, real-time tasks are scheduled first to ensure that users can obtain results in a timely manner.

[0079] Scheduling mechanism: Real-time queues use priority scheduling, usually First Input First Output (FIFO), but the system will dynamically adjust more urgent tasks to ensure that high-priority tasks can be processed immediately.

[0080] Resource allocation: This queue will prioritize system resources to ensure fast response.

[0081] For example, during a court trial, judges need to generate transcripts, extract key points, or analyze controversial issues in real time. These tasks are crucial to the trial's progress and must be completed in a very short time. Therefore, these tasks are prioritized within the real-time model queue.

[0082] (2) Asynchronous model priority queue (also called asynchronous queue):

[0083] Asynchronous priority queues are used to process tasks that have low response time requirements and can be processed later (called asynchronous tasks). These tasks do not require immediate results and can wait for system resources to be available before processing. Tasks in asynchronous queues have the following characteristics:

[0084] Low timeliness requirements: Asynchronous tasks do not need to be completed immediately and can take minutes, hours, or even longer. These tasks are suitable for background batch processing or non-urgent business processes.

[0085] Lower priority: The system will give priority to processing tasks in the real-time queue. Asynchronous tasks will only be processed when tasks in the real-time queue are completed or system resources are sufficient.

[0086] Scheduling mechanism: Asynchronous queues use batch scheduling and dynamically schedule tasks based on system load and resource usage. Low-priority tasks are deferred when resources are scarce.

[0087] Resource allocation: Asynchronous tasks do not compete with real-time tasks for system resources. Resources are allocated to asynchronous tasks only when the system load is low.

[0088] For example, batch processing of historical case data, periodic generation of document summaries, or batch classification of cases can be performed in the background and do not require immediate results. Therefore, such tasks can be placed in the asynchronous model priority queue.

[0089] The differences between the real-time model priority queue and the asynchronous model priority queue are as follows:

[0090] a. Timeliness

[0091] Real-time model priority queue: Tasks require immediate response and real-time results, typically within seconds or minutes. Timeliness is the most critical requirement.

[0092] Asynchronous model priority queue: Task processing can be delayed, is not time-sensitive, and allows background batch processing, which may be completed within a few minutes to a few hours.

[0093] b. Priority

[0094] Real-time model priority queue: The system allocates more computing resources to real-time tasks with higher priority to ensure fast execution. These tasks often directly impact the user's immediate operational experience or business processes.

[0095] Asynchronous model priority queue: The priority is low, and the task will not preempt the resources of the real-time task. Resources will be allocated for processing only when the system is idle.

[0096] c. Scheduling strategy

[0097] Real-time model priority queue: A first-come, first-served (FIFO) scheduling policy is adopted. However, based on the urgency of the task and the system load, the system may dynamically adjust the execution order of tasks to give priority to the most urgent tasks.

[0098] Asynchronous model priority queue: adopts a batch scheduling strategy, tasks will be processed in batches when resources are available, and the task execution order is relatively flexible, usually arranged according to the task creation time and system load.

[0099] d. System resource utilization

[0100] Real-time model priority queue: Tasks will try their best to obtain the system's computing resources, and system resources are allocated to real-time tasks first to ensure that tasks can be completed in a timely manner.

[0101] Asynchronous model priority queue: Tasks do not compete for system resources and are usually processed only when the system load is light and resources are idle. It is suitable for batch execution of resource-intensive tasks.

[0102] e. Applicable scenarios

[0103] Real-time model priority queue: Suitable for business scenarios that require rapid response, such as online judgment document generation, real-time case analysis, and other operations that require high immediacy.

[0104] Asynchronous model priority queue: Suitable for scenarios where backend batch processing and no rush to obtain results are required, such as historical case classification and data archiving.

[0105] In an exemplary embodiment, the scheduler 125 is configured to obtain the task from the real-time model priority queue 124 and / or the asynchronous model priority queue 123 , and send the case information of the task and the task configuration parameters to the model service layer 130 .

[0106] The scheduler plays the role of coordinator for task allocation and execution in the system architecture. It is responsible for obtaining tasks from the real-time model priority queue and the asynchronous model priority queue and distributing them to the corresponding model service for inference. The following is a detailed process for how the scheduler obtains model tasks from these two queues:

[0107] (1) Scheduler's task priority strategy

[0108] The scheduler usually uses a priority scheduling strategy to ensure that real-time tasks are processed first, while also processing asynchronous tasks as system resources allow. The following is the priority order in which the scheduler processes tasks:

[0109] Prioritize the real-time model priority queue: The scheduler first checks the tasks in the real-time model priority queue, as these tasks are often time-sensitive and must be processed immediately. The scheduler dynamically schedules tasks based on their priority, ensuring that the highest-priority real-time tasks are processed first.

[0110] Asynchronous tasks are processed when resources are idle: When tasks in the real-time queue are completed or system resources are sufficient (such as idle GPUs (Graphics Processing Units) and CPUs (Central Processing Units)), the scheduler extracts tasks from the asynchronous queue and processes them. Asynchronous tasks do not require processing, and the scheduler can schedule them when the system load is low.

[0111] (2) How the scheduler obtains tasks from the real-time model priority queue

[0112] a. Checking and processing real-time tasks: The scheduler periodically or continuously checks the real-time model priority queue and obtains and processes tasks according to the following steps:

[0113] Priority Check: The scheduler checks the priority of tasks in the queue, typically using a first-come, first-served (FIFO) policy, but the system may also dynamically adjust task priorities. For example, urgent tasks (such as generating transcripts for court trials) are directly picked up and processed first.

[0114] Task extraction: Based on the priority of the tasks in the queue, the scheduler will select the highest priority task for extraction and allocate it to the appropriate computing resources for inference processing.

[0115] Task execution: The scheduler assigns tasks to the corresponding model service. The model then performs inference based on the task's prompt word and context (including case information or processed case information), generating an inference result. Once processing is complete, the scheduler returns the inference result to upper-layer applications (such as the preprocessor and model application layer).

[0116] Load balancing: In scenarios with multiple real-time tasks, the scheduler may consider load balancing to ensure the proper allocation of system resources and avoid overuse of certain models or hardware resources, which could lead to system performance degradation.

[0117] b. Real-time Task Timeout Handling: If a real-time task takes longer than expected (for example, if model inference takes a long time), the scheduler can handle the timeout based on the task's deadline. Timed-out tasks may be rescheduled, or the system may return a default result, re-prioritize them, or simply discard them.

[0118] (3) How does the scheduler obtain tasks from the asynchronous model priority queue?

[0119] a. Task extraction when resources are idle: The scheduler obtains tasks from the asynchronous model priority queue only when there are no high-priority real-time tasks in the system or when system resources are idle. The specific steps are as follows:

[0120] Resource detection: The scheduler periodically checks the idle status of the model service. If an idle model service instance is found, the scheduler switches to the asynchronous model priority queue for task scheduling.

[0121] Adjust execution order as needed: Asynchronous tasks are usually not strictly processed on a first-come, first-served basis. The scheduler may adjust the execution order of tasks based on the urgency of the tasks, the current system load, etc. For example, when there is a serious backlog of tasks, the system may temporarily postpone some low-priority tasks.

[0122] Task Execution and Feedback: Similar to real-time tasks, after asynchronous tasks are assigned resources, the corresponding model performs inference processing. Once processing is complete, the scheduler returns the results to the upper application layer. Results can be returned when the system is idle, so users do not need to obtain results immediately.

[0123] b. Task retry and downgrade processing: For tasks in the asynchronous queue, the scheduler may retry or downgrade tasks that have not been processed for a long time:

[0124] Task retry: If a task cannot be processed in time due to some reasons (such as shortage of service resources), the scheduler can push the task back to the asynchronous queue and wait for the next scheduling opportunity.

[0125] Demotion processing: If there is a large backlog of tasks, the scheduler can demotion low-priority asynchronous tasks, delay their execution, or even cancel some unnecessary tasks.

[0126] By deploying asynchronous model priority queues and real-time model priority queues in the model application layer, the disclosed embodiment enables the preprocessor to determine the allocation of real-time tasks and asynchronous tasks to the corresponding real-time model priority queues and asynchronous model priority queues, respectively, based on the task priority and timeliness configuration obtained from the prompt word manager. This allows the scheduler to prioritize real-time tasks, ensuring a timely response to real-time tasks and avoiding delays in judicial trial progress. Furthermore, processing asynchronous tasks when resources are sufficient can improve the rationality of task processing and reduce resource usage.

[0127] In the exemplary embodiment, continue to refer to Figure 2 The model service layer 130 includes a unified standard model service gateway 131.

[0128] The scheduler 125 is also used to obtain the task from the real-time model priority queue 124 and / or the asynchronous model priority queue 123, use the case information as model input data, and encapsulate the model input data and the task configuration parameters into a remote procedure call (RPC) request packet, and pass the remote procedure call request packet to the unified standard model service gateway 131 through the remote procedure call protocol.

[0129] In an exemplary embodiment, when the type of the service model determined according to the degree of matching is a generative model, the prompt word manager 122 is also used to provide the preprocessor 121 with a prompt word template corresponding to the business scenario information, and model inference configuration options of the generative model, wherein the model inference configuration options include a generated content length limit and / or randomness control parameters.

[0130] In the embodiment of the present disclosure, the content length limit is used to guide the length of the text generated by the generative model to not exceed a certain limit. The specific limit can be set according to the actual scenario, and the present disclosure does not limit its value.

[0131] In an exemplary embodiment, the scheduler 125 is further configured to send the prompt word template and the model reasoning configuration options of the task to the model service layer 130. The task configuration parameters further include the prompt word template and the model reasoning configuration options.

[0132] In an exemplary embodiment, in the system architecture, the configuration that the preprocessor obtains from the prompt word manager is primarily related to the prompt words required for model inference and related task configuration. Specifically, the prompt word manager provides the preprocessor with one or more of the following information:

[0133] (1) Prompt word template.

[0134] The prompt word manager provides corresponding prompt word templates based on business scenario IDs (such as appeal analysis, dispute focus extraction, and document generation). These prompt word templates are designed to guide the AI ​​model to produce accurate output that meets scenario requirements. Prompt words can be combined with the judge's thought process to help the model better understand the task context.

[0135] (2) Model name (model ID)

[0136] Different business scenarios may require different big models (i.e., service models) for reasoning. The prompt word manager is configured with the big model name (model ID) used for each business scenario.

[0137] (3) Model parameters and strategy configuration

[0138] The prompt manager provides configuration options related to model inference (i.e., model inference configuration options), such as generation strategy (length limit for generated content, randomness control, etc.), inference temperature, Top-K, Top-P (also known as kernel sampling), and other generation model parameters. Inference temperature, Top-K, and Top-P are also called randomness control parameters and are used to control the randomness of model output.

[0139] The model's inference process primarily generates outputs based on the probability distribution of the generative model. During inference, the model generates the entire sentence step by step by predicting the next most likely word (or subword). This process is influenced by several parameters, such as temperature, Top-K, and Top-P.

[0140] In generative tasks, the inference temperature controls the smoothness of the output distribution. Higher temperatures lead to more randomness, while lower temperatures make the output more deterministic. Choosing the right temperature can help regulate the diversity and quality of generated results.

[0141] The Top-K and Top-P parameters control the output of the generative model. Top-K sampling limits the number of candidate words generated each time, while Top-P sampling (cumulative probability threshold) selects words whose cumulative probability exceeds p (a real number greater than 0 and less than 1). Appropriately setting these parameters can help improve the quality and diversity of generated text.

[0142] (4) Task priority and timeliness configuration

[0143] The prompt manager configures different task priorities based on the system's real-time asynchronous requirements. High-priority real-time tasks require the preprocessor to quickly process them and push them to the real-time model queue (i.e., the real-time model priority queue). Asynchronous tasks can be delayed and placed in the asynchronous priority queue (i.e., the asynchronous model priority queue).

[0144] In the system architecture, after the preprocessor obtains the configuration of the prompt word manager (i.e., task configuration parameters), it will assign model tasks to different queues based on the timeliness and priority of the tasks, namely the real-time model priority queue and the asynchronous model priority queue. The following is the assignment process:

[0145] (1) Judgment of timeliness

[0146] The preprocessor first determines the timeliness of the task based on the configuration provided by the prompt word manager. It can be divided into the following two tasks:

[0147] Real-time tasks require immediate processing and rapid return of results. For example, analyzing disputed points during a court hearing. These tasks typically have a high priority and strict response time requirements.

[0148] Asynchronous tasks: These tasks can be processed later and do not require immediate feedback. Examples include batch processing of historical case data and non-urgent document generation tasks. These tasks have a lower priority and can be placed in the asynchronous model's priority queue for processing.

[0149] (2) Priority judgment

[0150] The preprocessor also makes a priority judgment based on the specific configuration of the task (the preset priority level of the task).

[0151] High-priority tasks, such as urgent cases in court or tasks requiring immediate user feedback, are typically placed directly into the real-time model priority queue and queued according to priority.

[0152] Low-priority tasks, such as backend batch processing of case history data, do not require immediate results. These tasks typically enter the asynchronous model priority queue.

[0153] (3) Task allocation rules

[0154] Based on the above two conditions, the preprocessor performs the following operations: Real-time model priority queue: If a task is determined to be real-time and has a high priority, the preprocessor assigns the task to the real-time model priority queue. Tasks in this queue are immediately inferred and processed, and the results are quickly returned.

[0155] Asynchronous model priority queue: If a task is determined to be non-urgent and can be processed later, the preprocessor assigns it to the asynchronous model priority queue. Tasks in this queue are processed gradually based on the system's idle state, and the results may be returned later.

[0156] (4) Task scheduling mechanism

[0157] The two types of queues have different task scheduling mechanisms. Real-time priority queues typically use a first-in, first-out (FIFO) scheduling approach, but dynamically adjust based on task priority to ensure that high-priority tasks are processed quickly. Asynchronous priority queues typically employ a batch processing strategy that maximizes resource utilization. Tasks in these queues are gradually scheduled for processing as system resources become available. If a significant backlog of tasks occurs, the system may downgrade low-priority tasks or even delay their processing until sufficient system resources are available.

[0158] In an exemplary embodiment, the scheduler 125 is also used to obtain the task from the real-time model priority queue 124 and / or the asynchronous model priority queue 123, and then use the case information, the prompt word template and the model reasoning configuration options as model input data, and encapsulate the model input data and the task configuration parameters into a remote procedure call request package, and pass the remote procedure call request package to the unified standard model service gateway 131 through the remote procedure call protocol.

[0159] In the disclosed embodiment, after the scheduler obtains a model task from a queue (either a real-time model priority queue or an asynchronous model priority queue), it interacts with the unified standard model service gateway to assign the task to a specific model for inference. The role of the unified standard model service gateway is to coordinate the calls of multiple models, provide a unified interface, shield the differences in the underlying model services, and ensure load balancing and resource management of the models. The following is a detailed flow of the interaction between the scheduler and the unified standard model service gateway:

[0160] (1) Task acquisition and preprocessing

[0161] After the scheduler obtains a task from the real-time model priority queue or the asynchronous model priority queue, it first checks the task's configuration, including:

[0162] Task type: such as classification task, generation task, entity extraction task, etc.

[0163] Task parameters: including prompt words, context information, and inference parameters (such as generation length (i.e., the length limit of generated content), temperature, Top-K, etc.).

[0164] Model requirements: Some tasks may require calling specific models, such as text classification models, entity recognition models (i.e., entity extraction models), or generative models.

[0165] Based on these parameters, the scheduler prepares the input data required for the task (i.e., model input data) and decides which type of model to call.

[0166] (2) Interacting with the unified standard model service gateway

[0167] The scheduler interacts with the unified standard model service gateway to pass tasks to specific models for processing. The specific interaction process is as follows:

[0168] a. Calling standardized interfaces

[0169] The unified standard model service gateway provides a unified API (Application Programming Interface) interface for the scheduler to call. The scheduler sends a task request through this standard interface, which can include the following:

[0170] Model type requested by the task: Based on the task configuration, the scheduler will specify which model needs to be called.

[0171] Input data (model input data): includes the task input text, context information, prompt words (i.e., prompt word templates), etc. All of this input data will be packaged into a request and passed to the model gateway (i.e., the unified standard model service gateway).

[0172] Model parameters: The scheduler passes necessary inference parameters according to task requirements, such as the maximum length of the model generation task (i.e., the length limit of the generated content) and randomness control (such as temperature, Top-K, Top-P, and other parameters).

[0173] b. Load balancing and model selection

[0174] One of the functions of the Unified Standard Model Service Gateway (sometimes referred to as "gateway") is load balancing. After receiving a request from the scheduler, the gateway dynamically selects an appropriate model instance to handle the task based on the current system load. This process may involve the following operations.

[0175] Model instance selection: The unified standard model serving gateway can maintain multiple instances of the same model (such as multiple classification model instances or generative model instances). The gateway selects an appropriate model instance to handle the task based on the current load and available resources of each instance.

[0176] Resource management: The gateway checks the system's computing resources (such as GPU, memory, CPU, etc.) to ensure that the resources allocated to the model instances are sufficient to perform the tasks while avoiding overloading some model instances.

[0177] c. gRPC (high-performance remote procedure call protocol) protocol call

[0178] In the actual system implementation, the communication between the scheduler and the unified standard model service gateway uses the gRPC (remote procedure call) protocol. gRPC is an efficient communication protocol that supports cross-language communication and can well handle large-scale, low-latency service requests. The following is a typical gRPC call process:

[0179] Build request: The scheduler encapsulates the task input data (i.e., model input data), inference parameters, etc. into a gRPC request package (i.e., remote procedure call request package).

[0180] Sending a request: Through the gRPC protocol, the scheduler sends the request (i.e., remote procedure call request packet) to the gRPC server of the unified standard model service gateway.

[0181] Waiting for response: After receiving the request, the gateway forwards the task to the appropriate model instance for inference. After the inference is completed, the gateway encapsulates the result into a response and returns it to the scheduler.

[0182] d. Model inference process

[0183] When a model instance performs an inference task, the unified standard model service gateway is responsible for detecting the execution status of the model and returning the results to the scheduler after the task is completed. The model inference process mainly includes the following steps:

[0184] Model execution: After receiving a task, a model instance performs the corresponding inference operation. For example, a classification model (such as a text classification model) classifies the input data, while a generation model (i.e., a generative model) generates corresponding text based on the prompt word and input data.

[0185] Detection and feedback: The gateway detects the running status of the model instance to ensure that the task can be completed successfully. If a task encounters an error or timeout, the gateway will feedback to the scheduler for further processing (such as retry or downgrade).

[0186] Result return: After the model inference is completed, the result will be returned to the gateway through the gRPC protocol, and the gateway will return the inference result to the scheduler.

[0187] In an exemplary embodiment, the model service layer 130 includes a unified standard model service gateway 131, and a deep learning model 132 and a generative model 133 are deployed in the model service layer 130. Each deep learning model and each generative model has at least one model instance running in the model service layer 130.

[0188] The unified standard model service gateway 131 is used to receive the case information of the task and the task configuration parameters from the scheduler 125, and obtain the running status information of the model instance; obtain the input text vector of the case information of the task and the task configuration parameters, and obtain the model capability vectors of different types of service models; obtain the matching degree between the task and different types of service models based on the input text vector and the model capability vector; determine the corresponding deep learning model 132 or generative model 133 based on the matching degree; determine the corresponding model instance from the determined deep learning model 132 or generative model 133 based on the model instance running status information; distribute the case information of the task and the task configuration parameters to the corresponding model instance in the determined deep learning model 132 or generative model 133 for processing.

[0189] In an exemplary embodiment, a suitable model is selected based on the service request text. This can be achieved by matching the characteristics of the text content with the applicable scenarios of each model. The service request text is converted into a vector representation, and q is used to represent the service vector, that is, the input text vector. For different service models (i=1, 2, ..., N) uses the corresponding model capability vector Indicates different service models. Where N is a positive integer greater than or equal to 1.

[0190] In the embodiment of the present disclosure, the business request text refers to the formatted text content of the case information and task configuration parameters (optionally, including the prompt word template). The natural language is converted into an input text vector q through the embedding layer.

[0191] In the disclosed embodiments, a model capability vector is used to indicate the types of tasks that a service model can handle. For example, each service model has its own descriptive information (e.g., training dataset information, all fine-tuning directions, supported languages, parameter sizes, and examples of tasks that the model excels at). This descriptive information is converted into a vector using an embedding layer, which is the model capability vector.

[0192] In some embodiments, the cosine similarity method is used to calculate the matching degree based on the model capability vector and the input text vector, and the final result is to select the service model with the highest matching degree. For example, the matching degree is calculated using the following formula to measure the matching degree between the business request (i.e., task) and the service model:

[0193] (3)

[0194] Based on the calculated matching degree, select the service model with the highest matching degree:

[0195] (4)

[0196] The above formula means to select the service model with the highest matching degree with the task from N service models. A service model is used as the target model for this task. is a positive integer greater than or equal to 1 and less than or equal to N.

[0197] In the above embodiment, the type of model serving the task is determined based on the degree of matching, such as whether to call a small parameter summary model or a large parameter simple inference model. Then, the model instance running status information of the model instance of the service model of this type can be considered to determine the model instance for the task. The model instance running status information may include the load factor and response time of the model instance. For example, assuming that there are Z model instances of the service model serving the task based on the degree of matching, then the following formula can be used to determine which model instance is responsible for processing the task from these Z model instances:

[0198] (5)

[0199] In the above formula, Represents the load factor of the z-th model instance. The load factor indicates the indicator of the current busyness of the model instance. Model instances with lower loads are preferred. represents the response time of the zth model instance. Model implementations with shorter response times are preferred to improve efficiency. z and are all positive integers greater than or equal to 1 and less than or equal to Z, express The largest model instance. Z is a positive integer greater than or equal to 1. The disclosed embodiment can comprehensively consider business requests and load conditions, prioritize scheduling the most appropriate model and model instance, and improve system throughput and response speed.

[0200] In the disclosed embodiment, each model instance has its own generation queue, and each model instance has an average single token generation time, and the queue length of the generation queue of each model instance at the current moment is The average single token generation time is used to determine the response time of the model instance at the current moment. The queue length varies over time, while the average single token generation time is a fixed value for each model instance (determined by averaging historical stress tests). The generation queue refers to the collection of elements to be generated, arranged in a certain order or rule, when the model generates sequences (such as text or code). Elements in this queue are tokens, the basic units of model processing. In sequence generation tasks, the generation queue determines the order and content of the generated sequences. The model generates tokens one by one according to the order in the queue until a termination condition is met (such as a specific end symbol or a preset maximum length is reached). The generation queue length refers to the number of tokens to be generated in the queue. This length can be fixed or dynamic, depending on the specific task, model structure, and training strategy. The single token generation time refers to the time it takes for the model to generate the next token from the moment it receives the current input token. This time includes all aspects of the model's information processing, reasoning, and response generation. In deep learning, tokens are the basic units of input text sequences. These units can be words, characters, subwords, or other data fragments.

[0201] In an exemplary embodiment, the task configuration parameters also include the maximum response time of the task. In the embodiment of the present disclosure, a corresponding maximum response time can be set for each task type. For example, the maximum response time for generating a summary task can be set to be less than the maximum response time for performing a complex reasoning task. The unified standard model service gateway is also used to determine the remaining execution time and deadline of the task based on the maximum response time of the task; determine the comprehensive priority of the task based on the task priority, remaining execution time and deadline of the task; and within the current scheduling cycle, select the corresponding model instance of the task in the determined deep learning model or generative model for processing based on the comprehensive priority of the task.

[0202] In the disclosed embodiment, after determining which model instance of a service model to assign a task to for processing, the model instance may receive multiple tasks simultaneously. Therefore, the priority of each task can be used to determine which assigned task the model instance processes first. Task priority is used to properly schedule tasks in a multi-system, multi-task environment, ensuring that high-priority tasks are executed first. This embodiment employs a hybrid priority scheduling algorithm and combines multiple systems for comprehensive scheduling.

[0203] In some embodiments, the hybrid priority scheduling algorithm includes the following.

[0204] Set tasks The static priority of (It is pre-configured. Each task has its own priority, which can be the task priority in the task configuration parameters). The remaining execution time is , the deadline is , then the task The overall priority Defined as:

[0205] (6)

[0206] in, 、 、 They are 、 、 The weight coefficients represent the relative importance of the static priority, the remaining execution time, and the deadline, respectively, and are all greater than or equal to 0 and less than or equal to 1. In the embodiment of the present disclosure, a deadline is automatically assigned to the task when the task is generated, that is, the task must be completed before the deadline. If the task does not receive a response after the deadline, it is considered that the task has an error and an exception is handled. Remaining execution time = deadline - current time. Different maximum response times are set for different tasks, and the current time plus the maximum response time is the deadline. "Comprehensive" here means taking into account both the global priority of the task and the urgency of the current task, ensuring both the limited execution of high-priority tasks and the execution of urgent tasks (tasks that are about to reach the deadline). In each scheduling cycle, the scheduler selects the task with the highest comprehensive priority for processing based on the comprehensive priority.

[0207] In an exemplary embodiment, for priority scheduling of multi-system tasks, the goal is to minimize the overall delay of the system or maximize the throughput, and the task scheduling strategy is obtained by optimizing the objective function. Assume that the task set is , N is the number of tasks, N is a positive integer greater than or equal to 1, then the task scheduling goal can be defined as:

[0208] (7)

[0209] That is, by maximizing the sum of the comprehensive priorities of the tasks, the task throughput is maximized. In other embodiments, a scheduling strategy that minimizes the delay and response time of all tasks can also be adopted.

[0210] The hybrid priority scheduling algorithm focuses on prioritizing tasks within a single system. Multi-system task prioritization describes the resource preemption algorithm between different systems. The entire service platform contains different scheduling systems, such as civil production systems, criminal production systems, demonstration environments, testing environments, and specialized judge service systems. These systems share the same hardware service platform, necessitating an algorithm to balance the use of hardware resources across these systems.

[0211] In the disclosed embodiments, the unified standard model service gateway is an intermediate layer component used to manage and coordinate multiple model services (or service models) in the AI ​​system architecture. It provides a unified interface, shielding the differences between different underlying model services and simplifying the interaction between upper-layer applications and models. Through the unified standard model service gateway, the scheduler and upper-layer applications do not need to directly handle the details of each model. Instead, they can conveniently call, manage, detect, and load balance multiple model services through this gateway. Specifically, the functions of the unified standard model service gateway include the following aspects:

[0212] (1) Unified API interface

[0213] The unified standard model service gateway uses a unified API interface, allowing different model services to be called in the same way. Regardless of whether the underlying model is a deep learning model or a generative model, the scheduler and upper-level applications only need to interact with the gateway through a standard interface, avoiding the complexity of interacting directly with each model service.

[0214] Standardized request format: Regardless of the model type, the gateway requires all requests to be encapsulated in a unified format, such as task input, inference parameters, context information, etc. The specific request format is defined by the gRPC proto protocol (a binary data serialization protocol).

[0215] Unified return format: The gateway will standardize the output results of different models and return them to the upper-layer application in a unified format, so that the output of different models is consistent when used.

[0216] (2) Load balancing

[0217] The unified standard model service gateway is responsible for load balancing among multiple model instances to ensure efficient utilization of system resources and rapid processing of tasks.

[0218] (3) Model service detection and health check

[0219] The unified standard model service gateway is responsible for monitoring the health of each model service to ensure system stability and high availability. The gateway regularly performs health checks on each model instance to ensure that the model can respond to requests properly. If a model instance fails, the gateway marks it as unavailable, automatically restarts the service instance, and transfers tasks to other available instances. If a model instance fails, the gateway can automatically shut it down and automatically restart or restore the model service, ensuring system stability and high availability.

[0220] (4) Protocol management and scalability

[0221] The unified standard model service gateway typically uses an efficient communication protocol (gRPC) to interact with each model instance, ensuring low latency and high throughput. The gateway also offers excellent scalability, allowing new model services to be easily added or updated as needed without requiring changes to upper-layer applications.

[0222] (5) Security and access control

[0223] The unified standard model service gateway can manage and control external access, ensuring that only authorized services or users can call model services. It mainly uses authentication and authorization to ensure that only authenticated users or services can access model services.

[0224] In an exemplary embodiment, the corresponding model instance in the determined deep learning model 132 or generative model 133 is used to process the case information according to the task configuration parameters, obtain the reasoning result of the task, and return the reasoning result to the unified standard model service gateway 131.

[0225] The unified standard model service gateway 131 is also used to encapsulate the inference result into a response and return it to the scheduler 125.

[0226] In some embodiments, after receiving the result returned by the gateway, the scheduler may perform the following operations:

[0227] Result processing: The scheduler further processes the results returned by the model, such as result formatting and logging.

[0228] Feedback to upper-layer applications: The scheduler sends the final inference results back to the upper-layer application module (such as the scenario application layer) through the global message queue service, and the upper-layer application displays or stores them according to business needs.

[0229] Task status update: The scheduler records the status of the task (such as success, failure, or timeout) and performs retries or error handling when necessary.

[0230] In some embodiments, the scheduler can also be used for exception handling and task retry. In the process of interacting with the unified standard model service gateway, the scheduler needs to handle possible exceptions, including:

[0231] Timeout handling: If the model inference time exceeds the preset time limit, the scheduler can perform timeout handling, such as retrying the task or transferring the task to another model instance.

[0232] Model fault handling: If a model instance fails or cannot process a task, the unified standard model service gateway will return an error message. The scheduler can adjust the task execution strategy based on the error message, such as selecting other model instances for retry or downgrade processing.

[0233] Load regulation: If the system load is too high, the scheduler may temporarily postpone the execution of some low-priority tasks, or put the tasks back into the queue (asynchronous model priority queue or real-time model priority queue) to wait for resources to recover.

[0234] In an exemplary embodiment, when the type of the service model determined according to the matching degree is a generative model, the task configuration parameters further include the prompt word template and the model reasoning configuration options.

[0235] refer to Figure 3 In an exemplary embodiment, the generative model 133 includes one or more of the following models:

[0236] Summary models (e.g. Figure 3 The small parameter summary model 1331 in the example is used to generate one or more summaries of, for example, prosecution opinions, defense opinions, and appeal opinions based on the case information, the prompt word template, and the model reasoning configuration options.

[0237] The first type of inference models (e.g. Figure 3 A large-parameter simple reasoning model 1332 is used to process the case information, the prompt word template, and the model reasoning configuration options, perform legal reasoning, and generate text. In some embodiments, this can be used to extract key points of dispute and provide legal clause application recommendations.

[0238] The second type of inference models (e.g. Figure 3 The large-parameter complex reasoning model 1333 in the example is used to process the case information, the prompt word template, and the model reasoning configuration options to generate legal documents, which may include one or more of a judgment and an opinion.

[0239] In the generative models provided by the embodiments of this disclosure, different models have different parameter scales and complexities, suitable for different generative tasks. The following is a detailed introduction to small-parameter summary models, larger-parameter simple inference models, and large-parameter complex inference models, as well as their specific functions.

[0240] A small-parameter summary model refers to a vertical domain model with fewer than 10B parameters (for illustration purposes only, not limitation) trained on samples from the legal field. Its specific function is to generate summaries of case materials within the AI-assisted trial system, including prosecution opinions, defense opinions, and appeals, enabling judges and legal practitioners to quickly understand key case information.

[0241] Large-parameter simple reasoning models refer to vertical domain models with a parameter count of 30B to 40B (for illustration only, not limitation), trained on samples from the legal field. They are specifically used to perform simple legal reasoning tasks and generate text of moderate complexity within AI-assisted trial systems. Examples include the following tasks.

[0242] Extraction of dispute focus: Generates the main dispute focus based on the case content and provides simple legal explanations to assist judges in analyzing the case.

[0243] Suggestions on applicable legal provisions: Based on the description of the case, provide suggestions on applicable legal provisions, generate legal basis related to the case, and help judges refer to relevant legal provisions.

[0244] Large-parameter complex reasoning models refer to vertical domain models with a parameter count of 70B or more (for illustrative purposes only, not limited to this category) trained on samples from the legal field. They are specifically designed to handle highly complex generative reasoning tasks within AI-assisted trial systems, particularly those involving multiple rounds of reasoning or the generation of lengthy documents. For example, complex legal document generation involves generating detailed and complex legal documents, including judgments and opinions. These long, structured legal documents are suitable for handling complex cases.

[0245] It should be noted that the "small parameters", "large parameters" and "large parameters" in the embodiments of the present disclosure are relative, that is, different generative models have different amounts of model parameters after training. According to the different numbers of these model parameters, they are divided into small parameter summary models, large parameter simple inference models and large parameter complex inference models.

[0246] The generative model in the disclosed embodiments learns the joint probability distribution P(x,y) of the data, i.e., the joint distribution between features x and labels y. This means that it can not only predict labels for given features, but also generate new data instances. Generative models attempt to model the process of data generation, so data can be directly sampled from the model.

[0247] In the embodiments of the present disclosure, any one or more of the following major generative models can be used to implement the aforementioned small parameter summary model, large parameter simple reasoning model, and large parameter complex reasoning model:

[0248] 1. Recurrent Neural Networks (RNNs) are characterized by their ability to process sequential data, such as text or time series data. Each time step in an RNN takes in the current input and the hidden state from the previous time step, and outputs a new hidden state and a predicted value. Through continuous iteration, RNNs excel at generating text.

[0249] 2. Long Short-Term Memory (LSTM) is an improved RNN specifically designed to address long-term dependencies. It uses gating mechanisms (forget gate, input gate, and output gate) to control the flow of information, enabling more efficient learning of long-term dependencies. When generating new samples, LSTM units can be iterated to generate new data samples that match the characteristics of the input data.

[0250] 3. The Transformer is a model based on the self-attention mechanism. It can process input sequences in parallel, thus offering advantages in training and inference speed. A Transformer model consists of multiple attention heads and multiple self-attention layers. When generating new samples, the model can be fed with initial values ​​and iterated to generate new data samples that match the characteristics of the input data.

[0251] 4. Generative Adversarial Networks (GANs) consist of a generator network and a discriminator network. The generator network is responsible for generating fake data samples, while the discriminator network is responsible for distinguishing real data from fake data. Through adversarial training, the generator continuously improves to deceive the discriminator, while the discriminator continuously improves to better distinguish between real and fake data. Through continuous iterative training, the adversarial relationship between the generator and the discriminator becomes increasingly intense, and eventually the generator is able to generate new samples that closely resemble real data.

[0252] 5. Autoregressive models are a type of generative model based on probability distribution modeling. They work by establishing a joint distribution of data and using conditional probabilities to generate sequential data. During training, the model optimizes parameters by maximizing the posterior probability of the observed data and the latent variables, enabling the model to generate new samples that match the characteristics of the input data.

[0253] 6. The variational autoencoder is a generative model based on probabilistic coding that combines the ideas of autoencoders and variational inference. It consists of an encoder network that maps input data to a probability distribution in a latent space, while the decoder network samples and generates data from this distribution.

[0254] The deep learning model in the embodiments of this disclosure is a special machine learning model that learns the feature representation of data through multiple layers of nonlinear transformations. Deep learning models can be applied to generative models such as generative adversarial networks and variational autoencoders.

[0255] Continue to refer Figure 3 In an exemplary embodiment, the deep learning model 132 includes one or more of the following models:

[0256] The detection model 1321 is used to detect image information in the case information.

[0257] The text classification model 1322 includes a case cause classification model, which is used to process the case information and determine the case cause category of the case.

[0258] The entity extraction model 1323 is used to extract entity information from the case information.

[0259] The deep learning models provided in the embodiments of this disclosure include a detection model, a text classification model, and an entity extraction model. Each of these models is responsible for a different task and utilizes a deep learning algorithm appropriate for that task. The following are the deep learning algorithms for these models (it is understood that these algorithms are not limited to the following; other applicable deep learning algorithms may also be used. This is provided for illustrative purposes only) and their specific roles in the AI-assisted trial system:

[0260] (1) Detection model

[0261] YOLO (You Only Look Once): A real-time detection algorithm that quickly identifies and locates objects, suitable for scenarios requiring efficient detection. This disclosed embodiment uses YOLO6 to train multiple detection models, specifically for detecting ID cards, vehicle registrations, driver's licenses, lawyer's licenses, household registration books, seals, and QR codes in case materials. The detection models detect and classify the case materials at the front end, enabling the use of corresponding recognition models for information extraction.

[0262] (2) Text classification model

[0263] BERT (Bidirectional Encoder Representations from Transformers): BERT is a text classification model that leverages a bidirectional Transformer architecture to learn the contextual meaning of text, making it suitable for complex legal text classification tasks. RoBERTa, an improved version of BERT, can more efficiently process large corpus data and is suitable for high-precision text classification tasks. In this disclosed embodiment, a case classification model was trained based on RoBERTa. By inputting case materials or case descriptions, it automatically classifies cases into 486 categories of civil and commercial cases.

[0264] (3) Entity Extraction Model

[0265] The UIE (Universal Information Extraction) model, based on a pre-trained language model, is suitable for sequence labeling and information extraction tasks. It can capture contextual information in text and accurately extract essential information from legal documents. In this disclosed embodiment, the UIE model based on Paddle is retrained with legal data to extract company abbreviations, authorized principal information, and identity information in attorney letters.

[0266] In practice, how the gateway determines which model to call depends on the task configuration parameters obtained from upstream by the scheduler and the gateway proxy configuration. The model to call is determined based on the task configuration parameters. The unified standard model service gateway then distributes the task to the corresponding model through the proxy based on the model ID parameter. The task configuration parameters, excluding the model ID, are passed to the model service via RPC methods in the proto protocol.

[0267] For example, in the task of generating a summary, the scheduler obtains the input context information and task configuration parameters from the asynchronous model priority queue and / or the real-time model priority queue. This context information includes the summary prompt word template (also known as the prompt word instruction) and the case information content, which is the complete and processed content. The task configuration parameters include the model ID and inference strategy parameters (i.e., the inference parameters mentioned above). The unified standard model service gateway matches the model ID to the corresponding model service address in the proxy and passes the task content request.

[0268] The specific information output by the model depends on the type of task and the model used. The following are the information output by the model for several scenarios:

[0269] The output of a text classification model is usually a classification label, which indicates the specific category to which the input text belongs. The output information includes:

[0270] Classification tags: such as case causes, sales contract disputes, marriage disputes, motor vehicle traffic accidents, agency contract disputes, labor disputes, financial loans, credit card disputes, etc.

[0271] Classification confidence: When the model outputs the classification results, it usually provides a confidence classification for each category, indicating the model's confidence in the classification, such as 0.95.

[0272] Detection models are typically used for structural analysis of document images, and the output includes detected document elements and their location information. The output information includes:

[0273] Detected document elements: such as paragraphs, headings, tables, signatures, seals, etc.

[0274] Position coordinates: The position information of each element in the image (usually the coordinates of the bounding box), which is used to locate the specific position of the element in the document.

[0275] Figure 4 The embodiment is a schematic diagram of detecting various document elements in an image of a business license using a detection model and locating them using a bounding box.

[0276] The output of the entity extraction model is the key entities identified and annotated from the input text, such as names, places, legal terms, etc. The output information includes:

[0277] Entity list: such as the name of the party, the name of the lawyer, the location of the incident, etc.

[0278] Entity type: The type of each entity, such as "person name", "address", "legal terms", etc.

[0279] Entity position: The position in the original text, used to mark the start and end of the entity.

[0280] Figure 5 It is a schematic diagram of the output of an entity extraction model provided by an embodiment of the present disclosure.

[0281] The output of the generative model is usually natural language text, which generates corresponding content based on the input prompt words and context, and outputs the case summary content. Figure 6 Taking the generation of a summary as an example, the following example is shown:

[0282] The input prompt template is "You are a judge. Your task is to briefly summarize the facts and reasons in the plaintiff's complaint. The results are output in the following three directions:

[0283] 1. Basic facts: <…>

[0284] 2. Focus of appeal: <…>

[0285] 3. Legal provisions: <…>

[0286] The output format is: "Summary:"

[0287] After being processed by the generative model, the output is:

[0288] “1. Basic Facts:

[0289] 1.XXXXXXXXXXXXXXXXXXXXX

[0290] 2.…

[0291] 2. Focus of appeals:

[0292] 1. XXXXXXXXXXXXXXXXXXXXX

[0293] 2. XXXXXXXXXXXXXXXXXXXXX

[0294] 3. Legal provisions:

[0295] 1. XXXXXXXXXXXXXXXXXXXXX

[0296] 2.…”

[0297] In an exemplary embodiment, when the type of the service model determined according to the matching degree is a generative model, the task configuration parameters further include the prompt word template and the model reasoning configuration options.

[0298] refer to Figure 7 In an exemplary embodiment, the model service layer 130 further includes a first coordinator 134 .

[0299] The first coordinator 134 is used to obtain the case information of the task, the task configuration parameters, the prompt word template and the model reasoning configuration options from the unified standard model service gateway 131; obtain the model instance running status information corresponding to the generative model 133; determine to adopt the corresponding generative model according to the task configuration parameters; determine the corresponding model instance from the determined generative model according to the model instance running status information corresponding to the determined generative model; distribute the case information of the task, the prompt word template and the model reasoning configuration options to the corresponding model instance in the determined generative model for processing, obtain the reasoning result, and return the reasoning result to the unified standard model service gateway 131.

[0300] In the disclosed embodiment, the first coordinator mainly schedules generative model tasks and allocates resources rationally. The core role of the first coordinator is task management and scheduling. Its main functions in the execution of the generative model are as follows:

[0301] Concurrency control: When multiple tasks are submitted at the same time, the coordinator is responsible for controlling the concurrency of the tasks to avoid system overload.

[0302] Load balancing: The coordinator detects the load of each model instance and ensures that tasks are distributed across multiple model instances based on the minimum connection algorithm to avoid overloading some model instances while other instances are idle.

[0303] Task retry: If a build task fails due to resource issues or model failure, the coordinator can reschedule the task execution or assign it to another model instance.

[0304] Continue to refer Figure 7 In an exemplary embodiment, the model service layer 130 further includes a second coordinator 135 .

[0305] The second coordinator 135 is used to obtain the case information of the task and the task configuration parameters from the unified standard model service gateway 131; obtain the model instance running status information corresponding to the deep learning model 132; determine the corresponding deep learning model according to the task configuration parameters; determine the corresponding model instance from the model instance of the determined deep learning model according to the model instance running status information corresponding to the determined deep learning model; distribute the case information of the task to the corresponding model instance in the determined deep learning model for processing, obtain the inference result, and return the inference result to the unified standard model service gateway 131.

[0306] In an exemplary embodiment, a second coordinator can be configured on the deep learning model side. Its primary function is load balancing and concurrency control. As business volume increases, even though deep learning models consume less resources and process faster than generative models, a second coordinator is still needed to control the load. This ensures that the business system remains stable in high-concurrency scenarios.

[0307] Figure 8 This is a schematic diagram of another artificial intelligence-assisted trial system provided by an embodiment of the present disclosure. Figure 8 The artificial intelligence assisted trial system provided by the embodiment is Figure 7 Compared with the embodiment, the model application layer 120 further includes a feedback information processing node 126 .

[0308] In the disclosed embodiments, model output information refers to the inference results generated and returned by a generative model or deep learning model after completing an inference task. This output information may have different forms and contents depending on the model type, task requirements, and application scenario. The feedback information processing node is responsible for further processing and formatting these model outputs to ensure that the output results can be correctly understood and used by upper-level systems or users.

[0309] Continue to refer Figure 8 In an exemplary embodiment, the artificial intelligence assisted trial system 10 also includes a global message queue service 150.

[0310] In the disclosed embodiments, the global message queue service is a mechanism for asynchronous communication, task scheduling, and data transfer between modules within the system. It ensures efficient message delivery between different components, services, and modules, coordinating the overall workflow of the system, and is particularly crucial when handling concurrent and asynchronous tasks. The following are the specific functions of the global message queue service:

[0311] (1) Asynchronous task management and scheduling

[0312] The core function of the global message queue service is to manage asynchronous tasks and ensure that tasks can be effectively scheduled and executed between different modules.

[0313] Long-running tasks: Certain generative tasks, such as generating complex legal documents, can take a long time to complete. A global message queue can asynchronously process these long-running tasks, avoiding blocking other system operations.

[0314] Notification after task execution: When a task is completed, the processing service sends the execution result back to the message queue, and the queue service then sends the notification to the upper-level application to ensure that the task status is promptly fed back to the user or system.

[0315] (2) Decoupling and asynchronous communication between modules

[0316] The global message queue service effectively decouples system modules, allowing them to communicate asynchronously without requiring them to directly rely on each other or wait for each other's responses. This improves system flexibility. Modules (such as the model application layer, generative model service, and feedback information processing nodes) no longer need to communicate synchronously; they can instead send tasks to the message queue and continue processing other tasks. This avoids direct dependencies between modules and improves the system's parallel processing capabilities.

[0317] Prevent system congestion: The message queue can temporarily store tasks or data, and schedule the next task after completing a task, preventing the entire system from being blocked due to delays or bottlenecks in a certain module.

[0318] Asynchronous communication scenarios:

[0319] Asynchronous transmission of model output: After the inference task of the generative model or deep learning model is completed, the inference result is asynchronously transmitted to the feedback information processing node through the message queue, avoiding direct call dependency between the inference service and the feedback information processing node.

[0320] System notification and result delivery: When a module completes a task, it can notify other modules through the global message queue, ensuring that the task status is synchronized across all parts of the system. For example, when a document generation task is completed, the document review module or user interface is notified.

[0321] (3) Task reliability and fault tolerance mechanism

[0322] The global message queue service can also provide task reliability guarantees for the system, ensuring that tasks will not be lost during the sending, execution and feedback process, and can be recovered even if a partial failure occurs in the system.

[0323] Reliability assurance methods:

[0324] Message persistence: The message queue can persist the sent tasks, ensuring that even if the system fails during processing, the task information will not be lost, and the task can continue to execute after the system is restored.

[0325] Task retry mechanism: When a task fails during execution, the message queue can resend the task to other model instances for processing through the built-in retry mechanism to ensure that the task is eventually processed successfully.

[0326] (4) Event-driven processing mechanism

[0327] The global message queue service supports event-driven task processing. When an event is triggered (such as task completion, model inference success, etc.), the message queue can push a message to the modules that subscribe to the event, triggering subsequent operations. Event-driven scenarios include:

[0328] Reasoning task completion notification: When a generative model or deep learning model completes a task, the message queue pushes the completion information to the feedback information processing node or user interface, triggering subsequent feedback processing or result display.

[0329] Case analysis progress update: When asynchronously batch processing cases, the message queue can update the progress of the task and regularly send update information to the user interface or task management module to ensure that users are informed of the progress of the task in a timely manner.

[0330] (5) System scalability and high availability

[0331] The global message queue service provides support for the scalability and high availability of the system. Through the message queue, the system can easily expand new modules or model instances without modifying the existing business logic.

[0332] Manifestation of scalability: Dynamic expansion of modules or model instances: When the system needs to add new model instances or modules, new services can be dynamically connected through the message queue. The message queue will automatically distribute tasks to the newly added modules without the need to reconstruct existing services.

[0333] High availability of the system: The message queue can improve the system's fault tolerance through asynchronous processing of tasks and load balancing, avoiding the impact of a module failure on the overall business process.

[0334] Continue to refer Figure 8 In an exemplary embodiment, the artificial intelligence assisted trial system 10 also includes a global real-time communication service 160.

[0335] In the disclosed embodiments, the global real-time communication service refers to the communication component within the system used to transmit and synchronize information in real time, ensuring efficient coordination and instant messaging between various modules and services. The following describes the specific implementation mechanism of the global real-time communication service, as well as the specific meaning of "global" in this context.

[0336] Implementation mechanism of global real-time communication service: The global real-time communication service enables low-latency, instant messaging between multiple components and modules within the system, ensuring that tasks, status, and data can be quickly delivered to the required services or users. It is achieved through the following technical means:

[0337] a. Real-time communication based on message queues: The global real-time communication service relies on message queue systems (such as RabbitMQ) to implement real-time communication between multiple modules. Message queues support high-throughput, low-latency message delivery, ensuring real-time performance. This includes:

[0338] Message publish-subscription model: Each module within the system (such as model services, preprocessing nodes, and feedback processing nodes) can publish and subscribe to related messages. When a module generates a message, it is published to a queue, and all subscribers (related modules or services) can immediately receive these messages, ensuring real-time message delivery.

[0339] High availability and redundancy mechanisms: Message queue systems usually have high availability and redundancy mechanisms to ensure that messages are not lost in the event of network failure or service downtime.

[0340] b. WebSocket real-time two-way communication: For scenarios involving user interaction (such as the front-end of an AI-assisted adjudication system), WebSocket can be used to achieve real-time two-way communication. WebSocket maintains a persistent connection between the server and client, allowing the server to proactively push messages to the client, providing faster message delivery. Examples include the following scenarios:

[0341] Real-time task progress push: For example, when a user initiates a case analysis task, the system can provide real-time feedback to the user on the progress of task execution (such as status updates of document generation) through a WebSocket connection.

[0342] Instant notifications and reminders: When document generation is completed or case information processing is completed, the system can immediately notify the front-end or user end through WebSocket, without the client needing to poll the server to obtain status.

[0343] c. Efficient inter-service communication based on gRPC: gRPC (Google Remote Procedure Call) is a commonly used and efficient communication protocol for real-time communication between services. gRPC supports multiplexed, low-latency inter-service calls over HTTP / 2, making it particularly suitable for real-time messaging within systems.

[0344] Efficient binary transmission: gRPC uses Protocol Buffers for serialization and can transmit data in binary form, greatly improving the speed and efficiency of communication.

[0345] Bidirectional streaming communication: gRPC supports bidirectional streaming communication, which means that services can continuously exchange information with each other in the same connection, making it suitable for real-time communication service scenarios.

[0346] "Global" in this context has two meanings:

[0347] a. System-wide: "Global" refers to all modules and services within the system. This means that every component, service, and model instance within the system can participate in the real-time communication network and communicate with each other instantly. Information can be transmitted across services and modules, and is not limited to a specific service or module.

[0348] Real-time coordination between modules: For example, the document generation module (i.e., calling the AI-assisted trial system provided by the embodiment of the present disclosure to generate documents), the case analysis module (i.e., calling the AI-assisted trial system provided by the embodiment of the present disclosure to analyze cases), and the feedback information processing node all exchange information with each other through the global real-time communication service to coordinate task status.

[0349] Seamless service connection: Different services, such as the generative model service and the deep learning model service, can also interact through real-time communication services. For example, the classification results of a text classification model can be immediately passed to the generative model service to generate personalized legal documents.

[0350] In an exemplary embodiment, the global real-time communication service 160 includes a real-time communication service for enabling interaction between the deep learning model 132 and the generative model 133 .

[0351] b. Global across scenarios: Global also means that this communication service supports unified messaging and synchronization across multiple business scenarios, not just specific business processes. Whether it's real-time case analysis, batch document generation, or even archiving historical case data, all of these tasks can be delivered and synchronized using the same global real-time communication service.

[0352] Support for multiple business scenarios: The real-time communication service covers not only case processing, but also multiple business modules such as data archiving and user management, ensuring that all scenarios in the system can receive relevant notifications and messages.

[0353] In the disclosed embodiments, the role of the feedback information processing node is to further process the model output to meet the system requirements or user needs. The specific processing process varies depending on the task type and application scenario. An example is as follows:

[0354] a. Result formatting and structuring: The raw data output by the model may be unprocessed or semi-structured. The feedback information processing node will format or structure this data to facilitate system or user use.

[0355] Text classification result formatting: Convert the output of a text classification model into standardized labels or readable text information. For example, converting a classification result with a "classification label of 1" into a sales contract dispute case that is understandable to a judge.

[0356] Entity extraction result formatting: Structuring the entities and their types output by the entity extraction model, such as converting them into JSON format, to facilitate subsequent system calls or displays. Figure 9 It is a schematic diagram of the formatting of entity extraction results provided by an embodiment of the present disclosure.

[0357] b. Error handling and correction: Model output may sometimes contain erroneous or incomplete information. Feedback processing nodes need to correct or filter these results. Common processing methods include:

[0358] Confidence screening: For text classification or entity extraction models, if the confidence of an output is lower than the set threshold, the feedback information processing node will discard or mark the result to prevent erroneous information from affecting subsequent processing.

[0359] Post-processing of generated text: Generative models may sometimes output lengthy or repetitive text. The feedback information processing node will trim, remove duplication, or make semantic adjustments to the text based on the rules.

[0360] For example, for the judgment documents output by the generative model, the feedback information processing node may check whether the legal provisions are invalid to ensure that the generated documents comply with the specifications of legal documents.

[0361] c. Logging and Traceability: Feedback processing nodes may also log the results of each model inference to facilitate subsequent tracking and debugging. Log content includes: model input, model output, processed output, generation time, etc.

[0362] Figure 10 This is a flow chart of an artificial intelligence-assisted trial method provided by an embodiment of the present disclosure. Figure 9 As shown, the method provided by the embodiment of the present disclosure may include the following steps.

[0363] In S110 , the case information and business scenario information of the task are obtained.

[0364] In S120 , task configuration parameters of the task are acquired according to the business scenario information.

[0365] In S130 , the model instance running status information is obtained.

[0366] In S140 , the case information of the task and the input text vector of the task configuration parameters are obtained, and the model capability vectors of different types of service models are obtained.

[0367] In S150, the matching degree between the task and different types of service models is obtained according to the input text vector and the model capability vector.

[0368] In S160, based on the matching degree and the running status information of the model instance, the case information of the task and the task configuration parameters are distributed to the corresponding model instance of the corresponding type of service model for processing to obtain the reasoning result of the task to assist in the trial of the corresponding case.

[0369] Figure 10 The other contents of the embodiment can be referred to the other embodiments above, which will not be described in detail here. The following two practical application scenarios illustrate the operation of the AI-assisted trial system and method provided in the above embodiment.

[0370] Scenario 1: Generating the Focus of Dispute: The Focus of Dispute module (i.e., generating the focus of dispute by invoking the aforementioned AI-assisted trial system) clearly identifies the focus of dispute by deeply and meticulously comparing the prosecution opinions, defense opinions, and courtroom debates of the parties involved. The user clicks the "Join Generation Queue" button (i.e., the task trigger control for generating the focus of dispute) in the "Case Analysis - Trial Points" module of the front-end AI-assisted trial interface to submit a request for the task of generating the focus of dispute. Specifically, the following steps may be performed:

[0371] 1. Request Receipt and Analysis: The Intelligent-SDK service receives the request and analyzes the required input, such as the prosecution opinion, defense opinion, and court transcript (if available). The Intelligent-SDK service requests this case element information from the elementization module.

[0372] 2. Generating Input Content: After obtaining the required case element information, the Intelligent-SDK service assembles the case element information into a request body and sends it to the preprocessor. The preprocessor obtains the prompt word template for generating the focus of dispute and the relevant task configuration from the prompt word manager configuration as task configuration parameters. It then combines the case element information with the prompt word template to form the complete input content (i.e., model input data).

[0373] 3. Task Queuing and Allocation: The preprocessor assembles the input (i.e., case information) and task configuration parameters into a new request body according to the proto protocol. Based on the task configuration parameters, the preprocessor passes the request to the real-time model priority queue. The scheduler retrieves tasks from the real-time model priority queue and sends them to the first coordinator via the unified standard model service gateway based on the task configuration parameters. The first coordinator uses the least connection load algorithm (i.e., the least connection algorithm) to allocate tasks to idle model instances.

[0374] 4. Model processing and feedback: After the model instance completes its task, it returns the result to the scheduler through the unified standard model service gateway. The scheduler sends the result to the feedback information processing node for post-processing to generate the final result.

[0375] 5. Result display: The results are returned to the Intelligent-SDK service through the global message queue service. The Intelligent-SDK service displays the results on the front-end page for the judge to view.

[0376] Scenario 2: Generate and confirm facts:

[0377] The fact determination module (i.e., determining the facts of the case by calling the aforementioned AI-assisted trial system) collects and organizes case materials, conducts a preliminary review of the case facts, and provides users with a clear case fact framework to facilitate subsequent in-depth review and adjustment. The specific process is as follows:

[0378] 1. Request Receipt and Material Collection: The user initiates a request to generate a fact determination in the fact determination module of the AI-assisted trial interface (i.e., triggers the corresponding task trigger control). The Intelligent-SDK service receives the request and analyzes and generates the required basic materials, including the complaint, evidence, cross-examination opinions, and trial transcripts. The Intelligent-SDK service obtains and aggregates these basic materials from the elementization module as case information.

[0379] 2. Information Sorting and Integration: The Intelligent-SDK service organizes the collected materials into structured input content and sends it to the preprocessor. Based on the configuration of the prompt word manager, the preprocessor obtains the prompt word template and related task configuration required for fact determination as task configuration parameters, and combines the case element information and prompt word template into complete input content.

[0380] 3. Task Queuing and Distribution: The preprocessor constructs a request body based on the input content and task configuration parameters according to the proto protocol and passes the request to the asynchronous model priority queue based on the task configuration parameters. The scheduler extracts tasks from the asynchronous model priority queue and sends them to the second coordinator via the unified standard model service gateway. The second coordinator selects an idle model instance based on load and assigns the task.

[0381] 4. Model Processing and Preliminary Fact Generation: The model instance generates a preliminary factual framework for the case based on the input content and returns the results to the scheduler via the unified standard model service gateway. The scheduler sends the results to the feedback information processing node for post-processing to ensure that the output meets the requirements.

[0382] 5. Results Presentation and Further Adjustment: The generated preliminary factual framework is returned to the Intelligent-SDK service via the global message queue service and displayed on the front-end page for user reference. Users can review, verify, and adjust the generated factual framework based on their professional knowledge to form the final and complete case factual content.

[0383] In any of the above embodiments, when the front end interacts with the scene application layer, and when the Intelligent-SDK service interacts with the preprocessor, it can be achieved through the http (Hypertext Transfer Protocol) protocol, and this disclosure does not limit this.

[0384] It should be noted that in the technical solution of the present disclosure, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and require notification to users and obtaining their consent or authorization. When applicable, user personal information is de-identified and / or anonymized and / or encrypted.

[0385] Figure 11 This is a network architecture diagram for implementing an artificial intelligence-assisted trial system provided by the embodiment of the present disclosure. Figure 11 , which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present disclosure.

[0386] like Figure 11 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0387] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. For example, the terminal device serves as the front end 140 in the aforementioned AI-assisted trial system, and the scenario application layer 110, model application layer 120, and model service layer 130 are implemented through server 105. Terminal devices 101, 102, and 103 can be various electronic devices with display screens and web browsing support, including but not limited to smartphones, tablet computers, laptops, desktop computers, wearable devices, virtual reality devices, smart homes, and the like.

[0388] It is understandable that the terminal mentioned in the embodiments of the present disclosure may be a computer device, and the computer device in the embodiments of the present disclosure includes but is not limited to a terminal or a server. In other words, the computer device may be a server or a terminal, or a system consisting of a server and a terminal. The terminal mentioned above may be an electronic device, including but not limited to a mobile phone, a tablet computer, a desktop computer, a laptop computer, a PDA, an in-vehicle device, an augmented reality / virtual reality (AR / VR) device, a helmet display, a smart TV, a wearable device, a smart speaker, a digital camera, a camera, and other mobile internet devices (MID) with network access capabilities, or terminals in scenarios such as trains, ships, and airplanes.

[0389] The server 105 may be a server that provides various services, such as a background management server that provides support for devices operated by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received requests and other data, and feed back the processing results to the terminal device.

[0390] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as basic cloud computing services such as big data and artificial intelligence platforms, etc. This disclosure does not impose any restrictions on this.

[0391] It should be understood that Figure 11 The number of terminal devices, networks and servers is merely illustrative. The server 105 may be a single entity server or may be composed of multiple servers. It may have any number of terminal devices, networks and servers according to actual needs.

[0392] Optionally, the data involved in the embodiments of the present disclosure may be stored in a computer device, or the data may be stored based on cloud storage technology, which is not limited here.

[0393] Figure 12 This is a schematic diagram of the structure of a computer device provided by an embodiment of the present disclosure. Figure 12 , Figure 12 This is a schematic diagram of the structure of a computer device provided by an embodiment of the present disclosure. Figure 12 As shown, the computer device in the embodiment of the present disclosure may include: one or more processors 1201, memory 1202, and input / output interface 1203. The processor 1201, memory 1202, and input / output interface 1203 are connected via a bus 1204. The memory 1202 is used to store computer programs, which include program instructions. The input / output interface 1203 is used to receive and output data, such as for data exchange between a host machine and the computer device, or for data exchange between virtual machines in the host machine. The processor 1201 is used to execute the program instructions stored in the memory 1202. The processor 1201 can perform each step of the method shown in any of the above embodiments.

[0394] The embodiments of the present disclosure provide a computer device including a processor, an input / output interface, and a memory. The processor obtains a computer program in the memory to execute the steps of the method shown in any of the above embodiments.

[0395] The embodiments of the present disclosure also provide a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by the processor and executed by the artificial intelligence assisted trial system provided by each step in any of the above embodiments. For details, please refer to the implementation methods provided by each step in any of the above embodiments, which will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments involved in the present disclosure, please refer to the description of the method embodiments of the present disclosure. As an example, the computer program can be deployed to be executed on one computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected by a communication network.

[0396] The present disclosure also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in any of the optional embodiments described above.

Claims

1. An artificial intelligence-assisted trial system, characterized in that: include: The scenario application layer is used to obtain task case information and business scenario information; The model application layer is used to obtain case information and business scenario information of the task from the scenario application layer, and obtain task configuration parameters of the task according to the business scenario information; A model service layer, wherein different types of service models are configured, and the same type of service model runs at least one model instance, the model service layer is used to obtain model instance running status information, the model instance running status information includes the load factor and response time of the model instance, and obtain the case information and the task configuration parameters of the task from the model application layer; obtain the input text vectors of the case information of the task and the task configuration parameters, obtain model capability vectors of different types of service models, the model capability vectors include training data set information of the service model; and obtain the matching degree between the task and the different types of service models based on the input text vectors and the model capability vectors; The service model serving the task is determined based on the matching degree, and based on the load factor and response time in the model instance running status information of the model instance of the service model serving the task, the case information of the task and the task configuration parameters are distributed to the corresponding model instance of the corresponding type of service model for processing to obtain the reasoning result of the task for assisting in the trial of the corresponding case.

2. The system according to claim 1, wherein When the task includes a question, the system is further configured to: convert the question into a question vector; obtain a legal text vector of laws, regulations, and / or cases; calculate the semantic similarity between the question vector and the legal text vector; use the laws, regulations, and / or cases, as well as the question, as graph nodes, obtain edge weights between the graph nodes based on the semantic similarity between the graph nodes and a preset logical relationship of the trial thinking, use the question as a starting node, and find a reasoning chain through graph search; and obtain the credibility of the reasoning chain; According to the semantic similarity between the question vector and the legal text vector, and the credibility of the reasoning chain, a priority index of the candidate legal text is obtained; and according to the priority index, the laws, regulations and / or cases corresponding to the question are determined.

3. The system according to claim 1, wherein: The model application layer includes a preprocessor, a prompt word manager, an asynchronous model priority queue, a real-time model priority queue and a scheduler; wherein, The preprocessor is used to obtain the case information and business scenario information of the task from the scenario application layer, and send the business scenario information to the prompt word manager; The prompt word manager is used to receive the business scenario information, determine the corresponding model identification, task priority and timeliness configuration according to the business scenario information, and send the determined model identification, task priority and timeliness configuration to the preprocessor, wherein the task configuration parameters include the model identification, task priority and timeliness configuration; The preprocessor is further configured to push the task to the asynchronous model priority queue or the real-time model priority queue according to the task priority and timeliness configuration; The scheduler is used to obtain the task from the real-time model priority queue and / or the asynchronous model priority queue, and send the case information of the task and the task configuration parameters to the model service layer.

4. The system according to claim 3, wherein: The model service layer includes a unified standard model service gateway; The scheduler is also used to obtain the task from the real-time model priority queue and / or the asynchronous model priority queue, use the case information as model input data, and encapsulate the model input data and the task configuration parameters into a remote procedure call request package, and pass the remote procedure call request package to the unified standard model service gateway through the remote procedure call protocol.

5. The system according to claim 3, wherein: When the type of the service model determined according to the matching degree is a generative model, the prompt word manager is further configured to provide the preprocessor with a prompt word template corresponding to the business scenario information and model inference configuration options of the generative model, wherein the model inference configuration options include a generated content length limit and / or a randomness control parameter; The scheduler is further configured to send the prompt word template of the task and the model reasoning configuration options to the model service layer; The task configuration parameters also include the prompt word template and the model reasoning configuration options.

6. The system according to claim 5, wherein: The model service layer includes a unified standard model service gateway; The scheduler is also used to obtain the task from the real-time model priority queue and / or the asynchronous model priority queue, use the case information, the prompt word template and the model reasoning configuration options as model input data, and encapsulate the model input data into a remote procedure call request package, and pass the remote procedure call request package to the unified standard model service gateway through the remote procedure call protocol.

7. The system according to claim 3, wherein: The model service layer includes a unified standard model service gateway, and deep learning models and generative models are deployed in the model service layer, and each deep learning model and each generative model runs at least one model instance in the model service layer; The unified standard model service gateway is used to receive the case information of the task and the task configuration parameters from the scheduler, and obtain the running status information of the model instance; Obtaining the case information of the task and the input text vector of the task configuration parameters, and obtaining model capability vectors of different types of service models; obtaining the matching degree between the task and the different types of service models based on the input text vector and the model capability vector; determining to adopt a corresponding deep learning model or generative model based on the matching degree; determining a corresponding model instance from the determined deep learning model or generative model based on the model instance running state information; distributing the case information of the task and the task configuration parameters to the corresponding model instance in the determined deep learning model or generative model for processing; The corresponding model instance in the determined deep learning model or generative model is used to process the case information according to the task configuration parameters, obtain the reasoning result of the task, and return the reasoning result to the unified standard model service gateway; The unified standard model service gateway is also used to encapsulate the inference result into a response and return it to the scheduler.

8. The system according to claim 7, wherein: The task configuration parameters also include the maximum response time of the task; The unified standard model service gateway is also used to determine the remaining execution time and deadline of the task based on the maximum response time of the task; determine the comprehensive priority of the task based on the task priority, remaining execution time and deadline of the task; within the current scheduling cycle, select the corresponding model instance of the task in the determined deep learning model or generative model for processing based on the comprehensive priority of the task.

9. The system according to claim 7, wherein: When the type of the service model determined according to the matching degree is a generative model, the task configuration parameters further include a prompt word template and a model reasoning configuration option; the model service layer further includes a first coordinator; The first coordinator is used to obtain the case information of the task, the matching degree, the prompt word template and the model reasoning configuration option from the unified standard model service gateway; Obtaining the model instance running status information corresponding to the generative model; determining to adopt the corresponding generative model according to the matching degree; determining the corresponding model instance from the determined generative model according to the model instance running status information corresponding to the determined generative model; distributing the case information of the task, the prompt word template and the model reasoning configuration options to the corresponding model instance in the determined generative model for processing, obtaining the reasoning result, and returning the reasoning result to the unified standard model service gateway.

10. The system according to claim 7, wherein: When the type of the service model determined according to the matching degree is a generative model, the task configuration parameters further include a prompt word template and a model reasoning configuration option; wherein the generative model includes one or more of the following models: A summary model, configured to generate a summary based on the case information, the prompt word template, and the model reasoning configuration options; The first type of reasoning model is used to process the case information, the prompt word template and the model reasoning configuration options, perform legal reasoning and generate text; The second type of reasoning model is used to process the case information, the prompt word template and the model reasoning configuration options to generate legal documents.

11. The system according to claim 7 or 9, characterized in that The model service layer also includes a second coordinator; The second coordinator is used to obtain the case information and the matching degree of the task from the unified standard model service gateway; obtain the model instance running status information corresponding to the deep learning model; determine the corresponding deep learning model based on the matching degree; determine the corresponding model instance from the model instance of the determined deep learning model based on the model instance running status information corresponding to the determined deep learning model; distribute the case information of the task to the corresponding model instance in the determined deep learning model for processing, obtain the inference result, and return the inference result to the unified standard model service gateway.

12. The system according to claim 7, wherein: The deep learning model includes one or more of the following models: A detection model, used to detect image information in the case information; A text classification model, including a case cause classification model, wherein the case cause classification model is used to process the case information and determine the case cause category of the case; The entity extraction model is used to extract entity information from the case information.

13. The system according to claim 7, wherein: Also includes: A real-time communication service is used to enable interaction between the deep learning model and the generative model.

14. The system of claim 1, wherein: Also includes: A front end, configured to display an AI-assisted trial interface, the AI-assisted trial interface including a task trigger control, and in response to a trigger operation on a corresponding task trigger control, sending a request for the task to the scenario application layer, the request including a case identifier and request parameters; The scenario application layer is used to determine the business scenario identifier according to the request parameter, and obtain corresponding case information according to the business scenario identifier and the case identifier, wherein the business scenario information includes the business scenario identifier.

15. The system according to claim 14, wherein: The scenario application layer includes a case filing module, a middle platform module, a factorization module and an intelligent software development toolkit; The case filing module is used to receive case materials and perform structured processing on the case materials to obtain structured case data; The elementization module is used to obtain the case materials from the case filing module, extract elements from the case materials, and obtain case element information; The middle platform module is used to receive the request from the front end and send the request to the intelligent software development toolkit; The intelligent software development toolkit is used to determine the business scenario identifier according to the request parameter, obtain corresponding case data from the case filing module according to the business scenario identifier and the case identifier, and / or obtain corresponding case element information from the elementization module; and sending the obtained case data and / or case element information to the model application layer; The case information includes the case data and / or the case element information.

16. An artificial intelligence-assisted trial method, characterized in that: include: Obtain task case information and business scenario information; Acquire task configuration parameters of the task according to the business scenario information; Acquire model instance running status information, wherein the model instance running status information includes a load factor and a response time of the model instance; Obtaining case information of the task and an input text vector of the task configuration parameters, and obtaining model capability vectors of different types of service models, wherein the model capability vectors include training data set information of the service model; Obtaining a matching degree between the task and different types of service models according to the input text vector and the model capability vector; The service model serving the task is determined based on the matching degree, and the case information and the task configuration parameters of the task are distributed to the corresponding model instance of the corresponding type of service model for processing based on the load factor and response time in the running status information of the model instance of the service model serving the task, so as to obtain the reasoning result of the task for assisting in the trial of the corresponding case.

17. A computer device, characterized in that: Includes processor, memory, input and output interfaces; The processor is connected to the memory and the input / output interface respectively, wherein the input / output interface is used to receive and output data, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to claim 16.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so as to enable a computer device having the processor to perform the method of claim 16 .

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to claim 16 is implemented.

Citation Information

Patent Citations

  • Law case deep reasoning method and system

    CN118861197A

  • Data analysis method and device based on large model and nonvolatile storage medium

    CN119004026A