Big model-based multi-modal contract management method and device and electronic equipment

By employing a large-scale, multimodal contract management approach and utilizing machine learning models to process multimodal input contract instructions, the inefficiency of traditional manual operations is resolved, achieving efficient and intelligent contract management and improving the accuracy and compliance of intent determination.

CN120765204BActive Publication Date: 2025-12-05BEIJING AGILESTAR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511269445.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-05
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing technologies for contract management rely on manual operations, which are inefficient and have a high error rate in determining user intent, making it difficult to meet the needs of modern enterprises for efficient, intelligent, and compliant contract management.

Method used

A multimodal contract management method based on a large model is adopted. By acquiring the instruction sequence of multimodal input from users, a machine learning model is used to calculate the user's initial intent and probability, which is mapped to the intent semantic space. The intent vectors of different input methods are fused to generate the user's purpose intent and automatically execute contract management operations.

Benefits of technology

It improves the accuracy of intent determination, achieves seamless connection and efficient collaboration in all aspects of contract management, reduces the risk of incorrect intent judgment, and enhances the efficiency and compliance of contract management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765204B_ABST
    Figure CN120765204B_ABST
Patent Text Reader

Abstract

The application discloses a large model-based multi-modal contract management method and device and electronic equipment, and relates to the technical field of large models. The embodiment of the application can allow a user to input an instruction sequence using multiple input methods, thereby avoiding the risk of incorrect intention determination caused by the user inputting an intention using only one input method, and calculating an intention and a corresponding probability through a machine learning model, and converting the intention and the corresponding probability into an instruction intention vector based on the intention and the corresponding probability. Finally, the final purpose intention can be calculated based on each instruction intention vector, different input methods can be used to input instructions, the instructions input by different input methods can be calculated as attention weights, and the accuracy of intention determination is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large model technology, and in particular to a method, apparatus and electronic device for multimodal contract management based on large models. Background Technology

[0002] With the acceleration of enterprise digital transformation, contract management, as a core component of enterprise operations, is facing both severe challenges and vast opportunities. Traditional contract management methods rely primarily on manual operations, which are inefficient and prone to errors, especially when dealing with large-scale contract data, where their limitations are particularly pronounced. Furthermore, the continuous updates to laws and regulations and the frequent changes in industry standards further increase the complexity of contract management, making it difficult for enterprises to quickly adapt to dynamically changing compliance requirements and business needs.

[0003] Furthermore, contract management spans the entire contract lifecycle, including drafting, review, signing, performance, and archiving, making it an indispensable and crucial component of daily business operations. Contracts are not only fundamental documents for commercial cooperation but also vital tools for mitigating legal risks. However, current methods of contract management relying on manual processes suffer from numerous problems, such as low efficiency in manual review, fragmented information storage, lack of process standardization, and weak risk identification capabilities. These issues not only impact operational efficiency but may also increase legal risks, thereby jeopardizing the company's compliance and legal security.

[0004] To address the aforementioned issues, a multimodal contract management solution based on a large model is needed to overcome the shortcomings of existing technologies, such as reliance on manual processes leading to low management efficiency and high error rates in determining user intent. Summary of the Invention

[0005] This application provides a multimodal contract management method, apparatus, and electronic device based on a large model to address the shortcomings of existing technologies that rely on inefficient contract management and have a high error rate in determining user intent.

[0006] To achieve the above technical objectives, this application proposes a multimodal contract management method based on a large model, including:

[0007] Obtain a first instruction sequence input by the user, wherein the first instruction sequence includes at least two original instructions input by the user in at least two input methods, and the at least two original instructions are adjacent in chronological order of input time, and the time interval between the adjacent original instructions is shorter than a preset time threshold.

[0008] Based on the content of the at least two original instructions, a first machine learning model is used to calculate at least one initial user intent represented by each original instruction and its corresponding probability.

[0009] Based on the probability of each user's initial intent corresponding to each original instruction, each original instruction and its corresponding user initial intent are converted into an instruction intent vector, which is then mapped to the first intent semantic space.

[0010] Using a first machine learning model, the user's purpose intent corresponding to the first instruction sequence is calculated based on each instruction intent vector mapped to the first intent semantic space.

[0011] Perform the operation indicated in the user's intended purpose on the target contract indicated in the user's intended purpose.

[0012] This application also provides a multimodal contract management device based on a large model, including:

[0013] The acquisition module is used to acquire a first instruction sequence input by the user, wherein the first instruction sequence includes at least two original instructions input by the user in at least two input methods, and the at least two original instructions are adjacent in chronological order of input time, and the time interval between the adjacent original instructions is shorter than a preset time threshold.

[0014] An initial intent calculation module is used to calculate, based on the content of the at least two original instructions, at least one user initial intent represented by each original instruction and the corresponding probability using a first machine learning model;

[0015] The mapping module is used to convert each original instruction and its corresponding user initial intent into an instruction intent vector based on the probability of each original instruction corresponding to each user's initial intent, so as to map it to the first intent semantic space.

[0016] The purpose intent calculation module is used to calculate the user purpose intent corresponding to the first instruction sequence based on each instruction intent vector mapped to the first intent semantic space using a first machine learning model.

[0017] An execution module is used to perform the operation indicated in the user's intention on the target contract indicated in the user's intention.

[0018] This application also provides an electronic device, including:

[0019] Memory, used to store programs;

[0020] A processor is configured to run the program stored in the memory to execute a multimodal contract management method based on a large model according to embodiments of this application.

[0021] This application also provides a computer-readable storage medium storing a computer program executable by a processor, wherein the program, when executed by the processor, implements the multimodal contract management method based on a large model as provided in this application.

[0022] According to embodiments of this application, a multimodal contract management method, apparatus, and electronic device based on a large model calculate at least one user initial intent and corresponding probability represented by each original instruction using a first machine learning model based on the content of at least two original instructions input in a multimodal manner within the acquired user instruction sequence. Based on the probability of each user initial intent corresponding to each original instruction, each original instruction and its corresponding user initial intent are converted into instruction intent vectors and mapped to a first intent semantic space. Using the first machine learning model, the user's purpose intent corresponding to the first instruction sequence is calculated based on each instruction intent vector mapped to the first intent semantic space. Then, operations can be performed on the target contract according to the user's purpose intent. Therefore, it allows users to input instruction sequences using multiple input methods, avoiding the risk of incorrect intent judgment caused by users using only one input method. Furthermore, by calculating the time intent and corresponding probability using a machine learning model and converting it into an instruction intent vector, the final purpose intent can be calculated based on each instruction intent vector. Instructions input using different input methods can be fused together as attention weights, greatly improving the accuracy of intent determination.

[0023] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0024] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0025] Figure 1 A flowchart of an embodiment of the multimodal contract management method based on a large model provided in this application;

[0026] Figure 2 A schematic diagram of the structure of an embodiment of the multimodal contract management device based on a large model provided in this application;

[0027] Figure 3 A schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation

[0028] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0029] Example 1

[0030] Contract management, a core component of business operations, spans the entire contract lifecycle, including drafting, review, signing, performance, and archiving. It is a crucial means of ensuring smooth business cooperation and effectively mitigating legal risks. However, traditional contract management methods rely heavily on manual operations, resulting in inefficiency, error-proneness, and non-standardized processes, especially when handling large volumes of contracts. Furthermore, with continuous updates to laws and regulations and frequent changes in industry standards, the complexity of contract management continues to increase, posing significant challenges for businesses in addressing dynamically evolving compliance requirements and business needs.

[0031] In recent years, the accelerated digital transformation of enterprises has driven technological upgrades in the field of contract management. Intelligentization and automation have become the main directions for the development of contract management, with the application of artificial intelligence technology being particularly crucial. In particular, breakthroughs in large language models have provided new technological support for contract management. Large language models possess powerful language understanding and generation capabilities, enabling them to handle contract content in complex contexts and support multimodal data input (such as text, voice, and images), making them suitable for tasks such as contract generation, review, and optimization. Simultaneously, the integration of technologies such as knowledge graphs, semantic retrieval, and intelligent recommendation has further enhanced the accuracy and efficiency of contract management systems.

[0032] Despite the significant progress made in the aforementioned technologies, numerous challenges remain in practical applications. For example, the efficient processing of multimodal data needs further optimization, the accuracy of intent recognition in complex legal contexts urgently needs improvement, and the dynamic updating and expansion capabilities of knowledge bases are still imperfect. These technological bottlenecks make it difficult for existing contract management systems to fully meet the needs of modern enterprises for efficient, intelligent, and compliant contract management.

[0033] Specifically, existing technologies have proposed solutions using intelligent techniques to manage contracts based on user-input instructions. For example, users can input their desired operations via voice or text, such as generating or reviewing a contract. The system can then execute the corresponding operation by recognizing the user's voice or text input. However, in this case, since users typically lack professional knowledge of contract drafting or review, their instructions are often vague. For instance, in the scenario of generating a contract, although the system can recognize that the user intends to generate a contract, it cannot determine the specific type or requirements based on such simple input. Therefore, existing technologies can only generate a simple contract framework for the user, who then manually modifies it. Consequently, existing solutions are not only inefficient but also significantly hinder effective contract management due to inaccurate judgment of user intent.

[0034] Therefore, according to embodiments of this application, a multimodal contract management method based on a large model is provided, such as... Figure 1 As shown, Figure 1 This is a schematic flowchart illustrating a multimodal contract management method based on a large model according to an embodiment of this application. Figure 1 As shown in the embodiments of this application, the multimodal contract management method based on a large model may include:

[0035] S101, Obtain the first instruction sequence input by the user.

[0036] In step S101, a sequence of instructions consisting of multiple instructions input by the user in various ways can be received. For example, in this embodiment, the user can first take a screenshot or photograph of an existing contract interface and input the screenshot or photograph as an image as the first instruction. Then, through voice input, a new contract is generated using the contract as a template. Finally, the user can also input text, such as information about Party A and Party B, as well as other content that the user wants to include in the contract. Thus, these three original instructions input in different ways can constitute an instruction sequence. That is, such an instruction sequence can include at least two original instructions input by the user in at least two input methods, and these original instructions are adjacent in chronological order, and the time interval between adjacent original instructions is shorter than a preset time threshold. Therefore, if two of the multiple instructions input by the user detected in step S101 are approximately 10 seconds apart, while the third instruction is 10 minutes apart, then it can be considered that the third instruction may be unrelated to the first two instructions, and therefore it can be excluded from the instruction sequence.

[0037] Specifically, after step S101, for each instruction in the instruction sequence obtained in step S101, the original instruction can be processed according to the instruction input method. For example, when the input method of the original instruction is an image, the text in the image is extracted to obtain extracted text; the extracted text is then semantically converted using a pre-trained second machine learning model to obtain semantic text as the content of the original instruction input as an image. When the input method of the original instruction is speech, the speech is converted to text to obtain converted text; the converted text is then semantically converted using a pre-trained second machine learning model to obtain semantic text as the content of the original instruction input as speech; and when the input method of the original instruction is text, the pre-trained second machine learning model is used to perform semantic conversion to obtain semantic text as the content of the original instruction input as text.

[0038] S102, based on the contents of at least two original instructions, use a first machine learning model to calculate at least one user initial intent represented by each original instruction and the corresponding probability.

[0039] In step S102, a machine learning model can be used to calculate the user's initial intent and its corresponding probability based on the content of the multiple original instructions obtained in step S101. For example, based on an instruction input as an image, at least one semantic meaning of the image can be calculated, and this semantic meaning can be used as the user's initial intent, with the probability of the semantic meaning obtained during the calculation used as the probability of the initial intent.

[0040] S103, based on the probability of each user's initial intent corresponding to each original instruction, convert each original instruction and its corresponding user's initial intent into an instruction intent vector, so as to map it to the first intent semantic space.

[0041] In step S103, based on the probabilities of each initial intent calculated in step S102, each original instruction in the first instruction sequence obtained in step S101 and the user's initial intent determined in step S102 can be converted into intent vectors to map user intents of different modalities into the same semantic space. For example, in this embodiment, a first machine learning model can be used to calculate the first-level intent classification value of each instruction intent vector, wherein the first-level intent classification value indicates the probability that each instruction intent vector belongs to at least one first-level intent category. Then, based on the first-level intent classification value, the first machine learning model can be used to further calculate the second-level intent classification value of each instruction intent vector, which indicates the probability that each instruction intent vector belongs to at least one second-level intent category; based on the first-level intent classification value and the second-level intent classification value, the first machine learning model can be used to further calculate the third-level intent classification value of each instruction intent vector, which indicates the probability that each instruction intent vector belongs to at least one third-level intent category; finally, the third-level intent classification values ​​can be sorted, and one of the third-level intent categories can be determined as the user's purpose intent corresponding to the first instruction sequence based on the sorting result. Specifically, in this embodiment, the first-level intent classification can categorize user input into main categories such as "contract generation," "contract review," "contract signing," and "contract archiving." The second-level intent classification can further refine the first-level intent; for example, "contract generation" may include "automatic contract generation," "contract template matching," and "clause completion." The third-level intent classification can further refine the second-level intent; for example, "contract template matching" may include "matching templates based on business type" and "matching templates based on industry standards." Through this multi-level intent classification, the management method according to this embodiment can more accurately determine the user's current operational goal, thereby providing services that better meet their needs. Furthermore, in this embodiment, the user's historical operation information can also be recorded to form a data loop for subsequent intent recognition and recommendation optimization.

[0042] S104, using the first machine learning model, calculate the user's purpose intent corresponding to the first instruction sequence based on the instruction intent vectors mapped to the first intent semantic space.

[0043] In step S104, the first machine learning model can be used to further calculate the user's intended purpose from the instruction intent vectors determined in step S103. For example, the instruction intent vectors can be combined into a predetermined multimodal input structure, and the instruction intent vectors combined with the multimodal input structure can be fused to generate a unified feature representation. Specifically, the probability of each user's initial intent corresponding to each original instruction can be used as the weight of each instruction intent vector, and an attention mechanism can be used to calculate the attention value of the original instructions for each modality; the instruction intent vectors can be integrated based on the attention value to generate an integrated semantic feature vector as the unified feature representation.

[0044] Then, the integrated semantic feature vector can be feature encoded to generate an encoded semantic feature vector representing the semantic relationship and contextual information between the original instructions; a fully connected layer is used to calculate the probability that the first instruction sequence belongs to at least two preset intent categories on the encoded semantic feature vector; the probability distribution formed by each probability is normalized to generate a classification value, where the classification value represents the probability of the category of the user's purpose intent corresponding to the first instruction sequence; the probabilities are sorted, and according to the sorting result, the category with the highest probability is determined as the category to which the user's purpose intent belongs.

[0045] S105, Perform the operation indicated in the user's purpose and intent on the target contract indicated in the user's purpose and intent.

[0046] In step S105, the determined operation can be performed on the contract indicated by the user's intended purpose as determined in step S104. Step S104 can generate a complete contract workflow from the sequence of instructions formed by the multimodal instructions input by the user in step S101, thereby achieving seamless connection and efficient collaboration among various stages of contract management through automated process design. For example, the contract with the user's intended purpose determined in step S104 can include a generation process, a review process, and a signing process. This allows for the automatic generation of a contract draft based on user input or a selected template, prompting the user to supplement missing clauses; then, it automatically checks for special instructions identified in the user's original input to examine key clauses, such as liability for breach of contract and dispute resolution mechanisms, and marks potential risk points; finally, it allows the user to complete the signing online, and automatically records the signing time and signatory. Furthermore, after the contract is signed, the management method according to this application embodiment can automatically archive the signed contract and create an index for easy subsequent querying and retrieval.

[0047] Furthermore, in contract management scenarios, switching between multiple contract task windows is often involved. In this embodiment, based on the intent determined in step S104 and the generated workflow, the window for which the user needs to perform the operation can be automatically switched, thereby avoiding manual switching and improving work efficiency. Additionally, in this embodiment, real-time prompts or suggestions can be provided to the user during operation, based on the intent determined in step S104.

[0048] Furthermore, in this embodiment, a data closed-loop mechanism can be introduced to record user actions and execution results during the operation in step S105. For example, historical operations can be recorded, including input content, intent recognition results, and execution actions. During the operation, suggestions or evaluations of the execution process from the user can be received, allowing for adjustments to subsequent service strategies based on feedback. Additionally, based on historical data, the system can periodically train and optimize the intent recognition model and knowledge retrieval model to improve recognition accuracy and service quality.

[0049] According to the multimodal contract management method based on a large model according to the embodiments of this application, the method calculates at least one user initial intent and corresponding probability represented by each original instruction using a first machine learning model based on the content of at least two original instructions input in a multimodal manner in the obtained user instruction sequence; based on the probability of each user initial intent corresponding to each original instruction, each original instruction and corresponding user initial intent are converted into instruction intent vectors to be mapped to a first intent semantic space; using the first machine learning model, the user purpose intent corresponding to the first instruction sequence is calculated based on each instruction intent vector mapped to the first intent semantic space; and then the operation can be performed on the target contract according to the user purpose intent. Therefore, it allows users to input instruction sequences using multiple input methods, avoiding the risk of incorrect intent judgment caused by users inputting intent using only one input method. Furthermore, the machine learning model calculates the time intent and corresponding probability, and converts them into instruction intent vectors. Finally, the final purpose intent can be calculated based on each instruction intent vector. Instructions input in different input methods can be fused and calculated as attention weights, which greatly improves the accuracy of intent determination.

[0050] Example 2

[0051] Figure 2 This is a schematic diagram of the structure of one embodiment of the multimodal contract management device based on a large model provided in this application. Figure 2 As shown in the figure, the terminal maintenance device provided in this application embodiment may include: an acquisition module 21, an initial intent calculation module 22, a mapping module 23, a destination intent calculation module 24, and an execution module 25.

[0052] The acquisition module 21 can be used to acquire the first instruction sequence input by the user.

[0053] The acquisition module 21 can receive a sequence of instructions composed of multiple instructions input by the user in various ways. For example, in this embodiment, the user can first take a screenshot or photo of an existing contract interface and input the screenshot or photo as an image as the first instruction. Then, through voice input, a new contract is generated using the contract as a template. Finally, the user can also input text, such as information about Party A and Party B, as well as other content that the user wants to include in the contract. Thus, these three original instructions input in different ways can constitute an instruction sequence. That is, such an instruction sequence can include at least two original instructions input by the user in at least two input methods, and these original instructions are adjacent in chronological order, and the time interval between adjacent original instructions is shorter than a preset time threshold. Therefore, if the acquisition module 21 detects that two of the multiple instructions input by the user are approximately 10 seconds apart, while the third instruction is 10 minutes apart, then it can be considered that the third instruction may be unrelated to the first two instructions, and therefore it can be excluded from the instruction sequence.

[0054] Specifically, the acquisition module 21 can further process the original instructions according to their input methods for each instruction in the acquired instruction sequence. For example, when the original instruction is input as an image, the text in the image is extracted to obtain extracted text; the extracted text is then semantically converted using a pre-trained second machine learning model to obtain semantic text as the content of the original instruction input as an image. When the original instruction is input as speech, the speech is converted to text to obtain converted text; the converted text is then semantically converted using a pre-trained second machine learning model to obtain semantic text as the content of the original instruction input as speech; and when the original instruction is input as text, the pre-trained second machine learning model is used for semantic conversion to obtain semantic text as the content of the original instruction input as text.

[0055] The initial intent calculation module 22 can be used to calculate, based on the content of at least two original instructions, at least one user initial intent represented by each original instruction and the corresponding probability using a first machine learning model.

[0056] The initial intent calculation module 22 can use a machine learning model to calculate the user's initial intent and its corresponding probability based on the content of multiple original instructions obtained by the acquisition module 21. For example, based on an instruction input in the form of an image, at least one semantic of the image can be calculated, and this semantic can be used as the user's initial intent, with the probability of the semantic obtained during the calculation used as the probability of the initial intent.

[0057] The mapping module 23 can be used to convert each original instruction and its corresponding user initial intent into an instruction intent vector based on the probability of each user initial intent corresponding to each original instruction, so as to map it to the first intent semantic space.

[0058] The mapping module 23 can convert each original instruction in the first instruction sequence obtained by the acquisition module 21 and the user's initial intent determined by the initial intent calculation module 22 into intent vectors based on the probabilities of each initial intent calculated by the initial intent calculation module 22, so as to map user intents of different modalities into the same semantic space. For example, in this embodiment, a first machine learning model can be used to calculate the first-level intent classification value of each instruction intent vector, wherein the first-level intent classification value indicates the probability that each instruction intent vector belongs to at least one first-level intent category. Then, based on the first-level intent classification value, the first machine learning model can be used to further calculate the second-level intent classification value of each instruction intent vector, which indicates the probability that each instruction intent vector belongs to at least one second-level intent category; based on the first-level intent classification value and the second-level intent classification value, the first machine learning model can be used to further calculate the third-level intent classification value of each instruction intent vector, which indicates the probability that each instruction intent vector belongs to at least one third-level intent category; finally, the third-level intent classification values ​​can be sorted, and one of the third-level intent categories can be determined as the user's purpose intent corresponding to the first instruction sequence based on the sorting result. Specifically, in this embodiment, the first-level intent classification can categorize user input into main categories such as "contract generation," "contract review," "contract signing," and "contract archiving." The second-level intent classification can further refine the first-level intent; for example, "contract generation" may include "automatic contract generation," "contract template matching," and "clause completion." The third-level intent classification can further refine the second-level intent; for example, "contract template matching" may include "matching templates based on business type" and "matching templates based on industry standards." Through this multi-level intent classification, the management method according to this embodiment can more accurately determine the user's current operational goal, thereby providing services that better meet their needs. Furthermore, in this embodiment, the user's historical operation information can also be recorded to form a data loop for subsequent intent recognition and recommendation optimization.

[0059] The purpose intent calculation module 24 can be used to calculate the user purpose intent corresponding to the first instruction sequence based on each instruction intent vector mapped to the first intent semantic space using the first machine learning model.

[0060] The purpose intent calculation module 24 can further calculate the user's purpose intent using the first machine learning model on the instruction intent vector determined by the mapping module 23. For example, the instruction intent vectors can be combined into a predetermined multimodal input structure, and the instruction intent vectors combined with the multimodal input structure can be fused to generate a unified feature representation. In particular, the probability of each user's initial intent corresponding to each original instruction can be used as the weight of each instruction intent vector, and an attention mechanism can be used to calculate the attention value of the original instructions of each modality; based on the attention value, the instruction intent vectors can be integrated to generate an integrated semantic feature vector as the unified feature representation.

[0061] Then, the integrated semantic feature vector can be feature encoded to generate an encoded semantic feature vector representing the semantic relationship and contextual information between the original instructions; a fully connected layer is used to calculate the probability that the first instruction sequence belongs to at least two preset intent categories on the encoded semantic feature vector; the probability distribution formed by each probability is normalized to generate a classification value, where the classification value represents the probability of the category of the user's purpose intent corresponding to the first instruction sequence; the probabilities are sorted, and according to the sorting result, the category with the highest probability is determined as the category to which the user's purpose intent belongs.

[0062] Execution module 25 can be used to perform the operations indicated in the user's intent on the target contract indicated in the user's intent.

[0063] The execution module 25 can perform the determined operations on the contract indicated by the purpose intent calculated by the purpose intent calculation module 24. The purpose intent calculation module 24 can generate a complete contract workflow from the multimodal instruction sequence input by the user through the acquisition module 21, thereby achieving seamless connection and efficient collaboration of all aspects of contract management through automated process design. For example, the user purpose intent contract determined by the purpose intent calculation module 24 can include a generation process, a review process, and a signing process. Based on the user's input or selected template, a contract draft can be automatically generated, prompting the user to supplement missing clauses. Then, the module automatically checks for special instructions identified in the user's original input to examine key clauses, such as liability for breach of contract and dispute resolution mechanisms, and marks potential risk points. Finally, the user can complete the signing online, and the signing time and signatory are automatically recorded. Furthermore, after the contract is signed, the management method according to this application embodiment can automatically archive the signed contract and create an index for easy subsequent querying and retrieval.

[0064] Furthermore, in contract management scenarios, switching between multiple contract task windows is often involved. In this embodiment, based on the purpose intent determined by the purpose intent calculation module 24, the window for the user to perform the operation can be automatically switched according to the generated workflow, thereby avoiding manual switching by the user and improving work efficiency. In addition, in this embodiment, real-time prompts or suggestions can be provided to the user during operation based on the purpose intent determined by the purpose intent calculation module 24.

[0065] Furthermore, in this embodiment, a data closed-loop mechanism can be introduced, whereby the execution module 25 records user actions and execution results during the execution process. For example, historical operations can be recorded, including input content, intent recognition results, and execution actions. During the operation, suggestions or evaluations of the execution process from the user can be received, allowing for adjustments to subsequent service strategies based on feedback. Additionally, based on historical data, the system can periodically train and optimize the intent recognition model and knowledge retrieval model to improve recognition accuracy and service quality.

[0066] According to the multimodal contract management method based on a large model according to the embodiments of this application, the method calculates at least one user initial intent and corresponding probability represented by each original instruction using a first machine learning model based on the content of at least two original instructions input in a multimodal manner in the obtained user instruction sequence; based on the probability of each user initial intent corresponding to each original instruction, each original instruction and corresponding user initial intent are converted into instruction intent vectors to be mapped to a first intent semantic space; using the first machine learning model, the user purpose intent corresponding to the first instruction sequence is calculated based on each instruction intent vector mapped to the first intent semantic space; and then the operation can be performed on the target contract according to the user purpose intent. Therefore, it allows users to input instruction sequences using multiple input methods, avoiding the risk of incorrect intent judgment caused by users inputting intent using only one input method. Furthermore, the machine learning model calculates the time intent and corresponding probability, and converts them into instruction intent vectors. Finally, the final purpose intent can be calculated based on each instruction intent vector. Instructions input in different input methods can be fused and calculated as attention weights, which greatly improves the accuracy of intent determination.

[0067] Example 3

[0068] The above describes the internal functions and structure of a multimodal contract management device, which can be implemented as an electronic device. Figure 3 A schematic diagram illustrating the structure of an embodiment of the electronic device provided in this application. Figure 3 As shown, the electronic device includes a memory 31 and a processor 32.

[0069] Memory 31 is used to store programs. In addition to the programs described above, memory 31 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0070] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0071] Processor 32 is not limited to a central processing unit (CPU), but may also be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. Processor 32 is coupled to memory 31 and executes the program stored in memory 31. When the program runs, it executes the multimodal contract management method of Embodiment 1 described above.

[0072] Furthermore, such as Figure 3 As shown, the electronic device may also include other components such as a communication component 33, a power supply component 34, an audio component 35, and a display 36. Figure 3 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 3 The components shown.

[0073] Communication component 33 is configured to facilitate wired or wireless communication between electronic devices and other devices. The electronic devices can access wireless networks based on communication standards, such as WiFi, 3G, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 33 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 33 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0074] Power supply component 34 provides power to various components of the electronic device. Power supply component 34 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.

[0075] Audio component 35 is configured to output and / or input audio signals. For example, audio component 35 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 31 or transmitted via communication component 33. In some embodiments, audio component 35 also includes a speaker for outputting audio signals.

[0076] Display 36 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can detect not only the boundaries of the touch or swipe action, but also the duration and pressure associated with the touch or swipe operation.

[0077] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A large model-based multi-modal contract management method, characterized in that, The method comprises: obtaining a first instruction sequence input by a user, wherein the first instruction sequence comprises at least two original instructions input by the user in at least two input modes, and the at least two original instructions are adjacent in time sequence in terms of input time, and the time interval of the original instructions adjacent in time is shorter than a preset time threshold, calculating at least one user initial intent and a corresponding probability represented by each original instruction using a first machine learning model according to the content of the at least two original instructions; converting each original instruction and the corresponding user initial intent into an instruction intent vector based on the probability of the user initial intent corresponding to each original instruction, so as to be mapped into a first intent semantic space; calculating a user purpose intent corresponding to the first instruction sequence based on the instruction intent vectors mapped into the first intent semantic space using the first machine learning model; performing an operation indicated in the user purpose intent on a target contract indicated in the user purpose intent according to the user purpose intent. 2.The large model-based multi-modal contract management method of claim 1, wherein, The calculation of the user purpose intent corresponding to the first instruction sequence based on the instruction intent vectors mapped into the first intent semantic space using the first machine learning model comprises: calculating a first-level intent classification value of each instruction intent vector using the first machine learning model, wherein the first-level intent classification value indicates the probability that each instruction intent vector belongs to at least one first-level intent classification; further calculating a second-level intent classification value of each instruction intent vector using the first machine learning model based on the first-level intent classification value, wherein the second-level intent classification value indicates the probability that each instruction intent vector belongs to at least one second-level intent classification; further calculating a third-level intent classification value of each instruction intent vector using the first machine learning model based on the first-level intent classification value and the second-level intent classification value, wherein the third-level intent classification value indicates the probability that each instruction intent vector belongs to at least one third-level intent classification; ranking the third-level intent classification values, and determining one of the third-level intent classifications as the user purpose intent corresponding to the first instruction sequence according to the ranking result. 3.The large model-based multi-modal contract management method of claim 1, wherein, After obtaining the first instruction sequence input by the user, the method further comprises: when the input mode of the original instruction is a picture, extracting text in the picture to obtain extracted text; performing semantic conversion on the extracted text using a second machine learning model pre-trained to obtain semantic text as the content of the original instruction input in the picture mode, when the input mode of the original instruction is speech, performing text conversion on the speech to obtain converted text; performing semantic conversion on the converted text using a second machine learning model pre-trained to obtain semantic text as the content of the original instruction input in the speech mode, when the input mode of the original instruction is text, performing semantic conversion using a second machine learning model pre-trained to obtain semantic text as the content of the original instruction input in the text mode. 4.The large model-based multi-modal contract management method of claim 1, wherein, The using the first machine learning model, based on the mapping of each instruction intention vector in the first intention semantic space, includes: combining the instruction intention vectors into a predetermined multi-modal input structure; performing feature fusion on the instruction intention vectors combined in the multi-modal input structure to generate a unified feature representation. 5.The large model-based multi-modal contract management method of claim 4, wherein, The feature fusion on the instruction intention vectors combined in the multi-modal input structure to generate a unified feature representation includes: using the probability of each original instruction corresponding to each user initial intention as the weight of the instruction intention vector, and using the attention mechanism to calculate the attention value of the original instruction of each modality; integrating the instruction intention vectors based on the attention value to generate an integrated semantic feature vector as the unified feature representation. 6.The large model-based multi-modal contract management method of claim 5, wherein, The using the first machine learning model, based on the mapping of each instruction intention vector in the first intention semantic space, includes: performing feature encoding on the integrated semantic feature vector to generate an encoded semantic feature vector representing the semantic relationship between each original instruction and the context information; using a fully connected layer to calculate the probability that the first instruction sequence belongs to at least two preset intention categories; normalizing the probability distribution formed by each probability to generate a classification value, wherein the classification value represents the probability of the category of the user purpose intention corresponding to the first instruction sequence; sorting the probabilities, and according to the sorting result, determining the category to which the probability highest classification belongs as the category to which the user purpose intention belongs. 7.A large model-based multi-modal contract management apparatus, characterized by, It includes: an acquisition module configured to acquire a first instruction sequence input by a user, wherein the first instruction sequence includes at least two original instructions input by the user in at least two input modes, and the at least two original instructions are adjacent in time sequence and the time interval between the adjacent original instructions is shorter than a preset time threshold, an initial intention calculation module configured to calculate at least one user initial intention and a corresponding probability represented by each original instruction using a first machine learning model according to the content of the at least two original instructions; a mapping module configured to convert each original instruction and the corresponding user initial intention into an instruction intention vector based on the probability of each original instruction corresponding to each user initial intention, and map it to a first intention semantic space; a purpose intention calculation module configured to calculate a user purpose intention corresponding to the first instruction sequence using a first machine learning model based on the mapping of each instruction intention vector in the first intention semantic space; an execution module configured to execute an operation indicated in the user purpose intention on a target contract indicated in the user purpose intention according to the user purpose intention.

8. An electronic device, comprising: It includes: a memory configured to store a program; a processor configured to run the program stored in the memory to execute the large model-based multi-modal contract management method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intention category identification method and device

    CN111027667A

  • Instruction generation method and device of artificial intelligence accelerator and electronic equipment

    CN116339746A