DDD architecture service code examination and evaluation method based on multi-expert LLM
Through the field-driven design architecture of multi-expert large language model collaboration, the accuracy and efficiency of complex business code evaluation in the tobacco industry is solved, and efficient and intelligent automated code review is achieved to ensure code quality and consistency.
Patent Information
- Application Number
- CN202510378632.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology relies on manual review or a single expert model for code evaluation, resulting in limited accuracy and comprehensiveness of the evaluation results of complex business codes in the tobacco industry, and is inefficient, prone to bias and inconsistency.
The domain-driven design (DDD) architecture is adopted with a multi-expert large language model (LLM) collaboration, and automated code review is carried out through the token polling and delivery mechanism, and combined with the domain-driven design architecture, analyzing business code changes and providing optimization suggestions.
It improves the accuracy and comprehensiveness of code review, reduces manual dependence, improves the stability and efficiency of review, ensures that the code complies with industry architecture design standards, and promptly detects and fixes high-risk vulnerabilities.
Smart Images

Figure CN120276960A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and software process quality assurance, and in particular to a method for reviewing and evaluating business code in a DDD architecture based on a multi-expert LLM. Background Art
[0002] Currently, code review in software development mainly relies on manual review, static code analysis tools, and single-expert model review. Manual review depends on the experience and knowledge of developers, but the review efficiency is low and it is easily affected by personal biases and review fatigue. Although static code analysis tools can automatically check some common code problems, they often cannot deeply understand the business logic, the design and implementation of domain models, especially for the complex business code in the domain-driven design (DDD) architecture in the tobacco industry, and cannot provide high-quality evaluation and feedback. The root cause of these problems is that the existing technology relies on manual review or single-expert models for evaluation. The single-expert model is prone to the problem of large model hallucinations, and the reviewed code may have biases and inconsistencies. The root cause of these problems is that the existing technology relies on manual review or single-expert models for evaluation, and cannot analyze the business logic and architecture design of the code by using the cooperation of multi-expert large models in combination with business domain knowledge, resulting in limited accuracy and comprehensiveness of the evaluation results and low review efficiency. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides an automated code review and evaluation method under the domain-driven design (DDD) architecture based on the cooperation of a multi-expert large language model (LLM), which is mainly applied to the automated review and optimization of complex business code changes in the tobacco industry. This method collaborates through a multi-expert large model, combines the domain-driven design architecture, automatically analyzes business code changes, identifies potential problems, and gives optimization suggestions and scores.
[0004] The technical solution of the present invention is as follows:
[0005] A method for reviewing and evaluating business code in a DDD architecture based on a multi-expert LLM, comprising:
[0006] Reviewing the business code submission records through the cooperation of multi-expert large models, setting specific prompt templates for the expert large models based on the provided change code details, code dependencies, business requirements and other restrictions; constructing a broadcast common area where the expert large models can read each other's speech records and effectively communicate; achieving mutual checking and complementing among different expert large models under the common goal.
[0007] Develop an expert large model's speaking token polling mechanism. Specifically, the expert large model is divided into a leading large model and subordinate large models. Only the large model that holds the token can speak in the broadcast public area. The token transfer between subordinate large models adopts the token polling transfer mechanism, and finally the token is transferred back to the leading large model.
[0008] Design a code structure detection tool based on the characteristics of the domain-driven design DDD architecture. Perform code dependency parsing on the code context, and the parsing results will be used as the basis for the large model expert group to conduct code reviews.
[0009] Among them,
[0010] The token polling transfer mechanism is as follows:
[0011] Set M0 to represent the leading large model, M1, M2,..., M n to represent subordinate large models, T to represent the speaking transfer token (Token), and R i to represent the response status of subordinate large model M i . Among them, R i =1 indicates that the review request can be processed. Among them, R i =0 indicates that the review request cannot be processed. C i represents the speaking content and score of subordinate large model Mi, and P(M i ) represents the processing priority of subordinate large model M i , and k is the maximum number of polling times.
[0012] At the initial token allocation, set T = M0, that is, the leading large model M0 holds the token.
[0013] When the leading large model M0 finishes speaking for the first time and the token is transferred to a subordinate large model, a random selection scheme is adopted. Randomly select the subordinate large model M j with the highest priority from the subordinate large models that have not spoken;
[0014] M j = argmax{P(M i )|M i ∈{M1,..., M n} and M i has not spoken}
[0015] If R j = 1, it indicates that M j can process the request at this time, then transfer the token T = M j ; if R j = 0, it indicates that M j cannot process the request at this time, then the polling count is incremented by 1, and continue to randomly select other subordinate large models that have not spoken; if the polling count is greater than the maximum polling count k, then M jRemoved from the large model expert group.
[0016] When all subordinate large models have spoken or there are no subordinate large models, the leading large model summarizes its scores C0, C1, C2, ..., C with all subordinate large models n , and calculates the final score, where g is the comprehensive scoring function:
[0017] C = g(C0, C1, C2, ..., C n ).
[0018] Furthermore,
[0019] After selecting the specific business code change record, start to perform the automated code review evaluation analysis of the multi-expert large model collaboration under the domain-driven design DDD architecture.
[0020] The specific steps are as follows:
[0021] S10: Business framework integrity detection
[0022] Execute the customized code structure detection tool applicable to the domain-driven design DDD architecture of the tobacco industry to perform framework integrity detection on the changed files, and store the results in the framework non-integrity detection record.
[0023] S20: Code dependency relationship parsing
[0024] Execute the code dependency parsing tool related to the changed code context to perform hierarchical dependency parsing, capture the upper and lower layer business codes related to the changed code, and store the relevant business dependencies of the changed code in the business code dependency record.
[0025] S30: Construction of multi-expert large model collaboration framework
[0026] Adopt the multi-expert large model collaboration framework, construct the large model code review expert group, confirm the leading large model and subordinate large models, issue the token for controlling the large model's speech to the leading large model, and formulate the review rules and the Prompt template for the large model, including role requirements, optimization requirements, and specific scoring rules.
[0027] S40: Final evaluation and alarm of the code review system
[0028] After the code review system receives the changed submitted code, summarize the code submission rating, that is, the discrimination is excellent, qualified, and unqualified. The submission records with the code submission rating result of unqualified will issue a system alarm to urge for secondary code review to facilitate timely code rectification and give priority to eliminating high-risk vulnerabilities.
[0029] Even further,
[0030] Leading large model code review evaluation
[0031] The leading large model conducts code review and evaluation according to the framework non-integrity detection records, business code dependency records, details of changed code, commit message of the submission, and the tendency of the project preference tool class, and gives a code score according to the scoring structure. After the review is completed, the code review and evaluation results and optimization suggestions are sent to the broadcast public area of the multi-model collaboration framework, and the token is passed to the subordinate large model through the token polling algorithm.
[0032] Speaker token polling algorithm
[0033] The expert token polling algorithm will detect the subordinate large model currently in the expert group for code review of the large model, randomly select a subordinate large model that has not spoken yet to check the online status of this subordinate large model, and send a request to check whether the basic requirements for code review can be replied in a timely manner. If this subordinate large model returns a response status code indicating that it can reply to the code review request in a timely manner, the speaking token will be officially passed to this subordinate large model; otherwise, the speaking token will be passed to the next subordinate large model. If one of the subordinate large models fails to respond correctly multiple times, it will be temporarily removed from the target subordinate large models of this round of speaker token polling algorithm.
[0034] Subordinate large model code review and evaluation
[0035] The subordinate large model obtains the relevant information mentioned above through the broadcast public area, analyzes potential problems and provides optimization suggestions for the provided changed code according to the prompt words, and gives a code score under the specified scoring framework. After sending the answer result to the broadcast public area of the multi-model collaboration framework, the token is passed to the next subordinate large model through the token polling algorithm. And so on, until the last subordinate large model conducts code review and evaluation and scores, and the token is passed to the leading large model.
[0036] Leading large model summary and final score
[0037] When the last subordinate large model completes the review, the final result is summarized by the leading large model, listing the problems and optimization suggestions existing in the submitted code for this change item by item. According to the scoring results of all expert large models, the final score result is summarized according to the expert scoring algorithm, and the final score result and evaluation suggestions are sent to the code review system.
[0038] The beneficial effects of the present invention are
[0039] The present invention creates an efficient, intelligent, and stable automated code review and evaluation method through the collaboration of multiple expert large models, token polling transmission, and specific scenario domain-driven design (DDD) architecture analysis. Compared with traditional manual code review or single large model evaluation methods, the present invention has significant advantages in terms of accuracy, stability, industry adaptability, efficiency improvement, security, etc., and can provide more reliable technical support for the business code review of the tobacco industry, and is widely applicable to the code review scenarios of domain-driven design architectures in other industries.
[0040] In the same submitted code group of the present invention, different leading large models can be tried as the leading large model, and there are differences in the number of final code review optimization suggestions. The currently mainstream DeepSeek-R1, Llama-3, and GLM-4 are selected as the leading large models for comparison respectively, and the number of leading large model code review optimization suggestions summarized in 4 code groups is shown. It is concluded that the large model with better theoretical performance as the leading large model has better effects and provides more optimization suggestions. The final result needs to be determined according to the context Token number supported by the large model.
[0041] The present invention effectively improves the accuracy and comprehensiveness of code review and reduces the manual dependence on code review evaluation. Through the multi-expert large model collaboration mechanism, environmental problems easily generated by a single large model are avoided, and the review stability is improved. Through the token polling mechanism of the large model expert group, the collaboration efficiency between large models is improved, ensuring the fair speech of each large model. The online status detection mechanism for subordinate large models can ensure that the token is passed to the currently available large model, avoiding the impact of individual large models on the overall review process, and effectively improving the business code review efficiency of the system. For the business code requirement scenario under the domain-driven model (DDD) architecture in the tobacco industry, the formulated code structure detection tool can ensure that the code complies with the industry's architecture design standards and business rules. By formulating quantifiable review code scoring rules, warnings are issued for unqualified code submissions, facilitating timely reminders to developers for manual rechecks, thereby ensuring the stability of the system while improving the system development efficiency. Brief Description of the Drawings
[0042] Figure 1 is a schematic diagram of the working process of the present invention;
[0043] Figure 2 is a schematic diagram of the number of optimization suggestions for leading large model code review;
[0044] Figure 3 is a schematic diagram of the architecture of the present invention. Detailed Embodiment
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0046] The present invention provides a method for reviewing and evaluating business code in a DDD architecture based on a multi-expert large language model (LLM), aiming to improve the software quality of business code, ensure that business programs strictly follow software development plans, enhance the maintainability and consistency of code, meet the specific requirements of the tobacco industry, and provide efficient support for subsequent development and maintenance.
[0047] The main contents of the present invention include:
[0048] ① Multi-expert large model collaboration mechanism
[0049] The review of business code submission records is carried out through the collaboration of multi-expert large models. Based on the provided details of changed code, code dependencies, and business requirements, specific prompt templates are set for the expert large models. A broadcast common area is constructed, and the expert large models can effectively read the speech records between each other and speak effectively. Different expert large models can check and make up for each other's deficiencies under the common goal, eliminate common "hallucination" problems of large models through multi-angle analysis and optimization, avoid biases and inconsistencies caused by a single large model, and finally form a code review result and review score that meet the final result.
[0050] ② Expert large model speech token polling mechanism
[0051] The present invention formulates an expert large model speech token polling mechanism for the collaboration of multi-expert large models. Specifically, the expert large models are divided into a leading large model and subordinate large models (there are multiple subordinate large models). Only the large model that holds the token can speak in the broadcast common area. The token transfer between subordinate large models adopts a token polling transfer mechanism, and finally the token is transferred back to the leading large model. The token polling transfer mechanism is as follows:
[0052] Assume that M0 represents the leading large model, M1, M2,..., M n represents the subordinate large models, T represents the speech transfer token (Token), and R i represents the response status of the subordinate large model M i (where R i = 1 indicates that the review request can be processed, and where R i = 0 indicates that the review request cannot be processed), and C i represents the subordinate large model Mi The speech content and score of P(M i ) indicates the processing priority of the subordinate large model M i , and k is the maximum polling count.
[0053] At the initial token allocation, let T = M0, that is, the leading large model M0 holds the token.
[0054] When the leading large model M0 finishes speaking for the first time and the token is passed to the subordinate large model, a random selection scheme is adopted. Randomly select the subordinate large model with the highest priority M j .
[0055] M j = argmax{P(M i )|M i ∈{M1,..., M n} and M i has not spoken}
[0056] If R j = 1, it indicates that M j can process the request at this time, then pass the token T = M j ; if R j = 0, it indicates that M j cannot process the request at this time, then the polling count is incremented by 1, and continue to randomly select other subordinate large models that have not spoken. If the polling count is greater than the maximum polling count k, then M j is removed from the large model expert group.
[0057] When all subordinate large models have spoken or there are no subordinate large models, the leading large model summarizes its own scores C0, C1, C2,..., C n with all subordinate large models, and calculates the final score (where g is the comprehensive scoring function):
[0058] C = g(C0, C1, C2,..., C n )
[0059] ③ Customize industry-specific business framework requirements (Domain-Driven Design architecture)
[0060] The present invention can design a code structure detection tool based on the characteristics of the Domain-Driven Design (DDD) architecture according to the specific requirements of the tobacco industry, and can perform code dependency parsing on the code context. The parsing result will be used as an important basis for the large model expert group to conduct code reviews.
[0061] Detailed steps:
[0062] After selecting the specific business code change record, start to execute the automated code review and evaluation analysis of the multi-expert large model collaboration under the domain-driven design (DDD) architecture. The specific steps are as follows:
[0063] S10: Business framework integrity detection
[0064] Execute the code structure detection tool customized for the domain-driven design (DDD) architecture of the tobacco industry to detect the framework integrity of the changed files and store the results in the framework non-integrity detection record.
[0065] S20: Code dependency relationship parsing
[0066] Execute the code dependency parsing tool related to the context of the changed code to perform hierarchical dependency parsing, capture the upper and lower layer business codes related to the changed code, and store the related business dependencies of the changed code in the business code dependency record.
[0067] S30: Construction of multi-expert large model collaboration framework
[0068] Adopt the multi-expert large model collaboration framework to construct a large model code review expert group, confirm the leading large model and subordinate large models, issue the token for controlling the large model's speech to the leading large model, and formulate the review rules and the prompt template for the large model, including role requirements, optimization requirements, and specific scoring rules.
[0069] S31: Code review and evaluation by the leading large model
[0070] The leading large model conducts code review and evaluation according to the content of the framework non-integrity detection record, business code dependency record, details of the changed code, commit message, project preference tool class tendency, etc., and gives a code score according to the scoring structure. After the review, send the code review and evaluation results and optimization suggestions to the broadcast common area of the multi-model collaboration framework, and pass the token to the subordinate large models through the token polling algorithm.
[0071] S32: Speech token polling algorithm
[0072] The expert token polling algorithm will detect the subordinate large models in the large model code review expert group, randomly select a large model that has not spoken yet to check the online status of the current subordinate large model, and send a request to check whether the basic requirements for code review can be replied in time. If the subordinate large model returns a response status code indicating that it can reply to the code review request in time, then officially pass the speech token to the subordinate large model. Otherwise, pass the speech token to the next subordinate large model. If one of the subordinate large models fails to respond correctly multiple times, temporarily exclude it from the target subordinate large models of this round of speech token polling algorithm.
[0073] S33: Subordinate large model code review and evaluation
[0074] The subordinate large model obtains information related to the previous text through the broadcast common area, analyzes potential problems in the provided changed code and provides optimization suggestions according to the prompt words, and gives a code score under the specified scoring framework. After sending the answer result to the broadcast common area of the multi-model collaboration framework, the token is passed to the next subordinate large model through the token polling algorithm. And so on, until the last subordinate large model conducts code review evaluation and scoring, and passes the token to the leading large model.
[0075] S34: Leading large model summary and final scoring
[0076] When the last subordinate large model completes the review, the leading large model summarizes the final results, itemizes the problems and optimization suggestions existing in the changed submitted code, summarizes the final score result according to the scoring results of all expert large models based on the expert scoring algorithm, and sends the submitted final score result and evaluation suggestions to the code review system.
[0077] S40: Final evaluation and alarm of the code review system
[0078] After receiving the changed submitted code, the code review system summarizes the code submission rating (the differentiation is excellent, qualified and unqualified). The submission records with the code submission rating result of unqualified will issue a system alarm to urge the project supervisor to conduct a secondary code review manually, so as to facilitate timely code rectification and give priority to eliminating high-risk vulnerabilities.
[0079] Through the method for evaluating business code review of the domain-driven design (DDD) architecture based on multiple expert large language models (LLMs) of the present invention, the business code of the domain-driven design architecture can be efficiently evaluated in the tobacco industry. The accuracy of the evaluation is improved through multi-angle analysis by multiple expert large models. The large models correct themselves and discuss with each other during the review process, forming an effective review result feedback mechanism to promote the optimization of business code.
[0080] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for reviewing and evaluating business code in a DDD architecture based on a multi-expert LLM, characterized in that: It includes: Review the business code submission records through the cooperation of multi-expert large models. Based on the provided details of changed code, code dependencies, and business requirements, etc., set specific prompt templates for the expert large models; construct a broadcast common area where the expert large models can read each other's speech records and communicate effectively; achieve mutual checking and filling of gaps among different expert large models under the common goal. Formulate a token polling mechanism for the expert large models to speak. Specifically, the expert large models are divided into a leading large model and subordinate large models. Only the large model holding the token can speak in the broadcast common area. The token transfer among subordinate large models adopts a token polling transfer mechanism, and finally the token is transferred back to the leading large model. Design a code structure detection tool based on the characteristics of the Domain-Driven Design (DDD) architecture, perform code dependency parsing on the code context, and the parsing results will be used as the basis for the code review by the expert large model group.
2. The method according to claim 1, characterized in that: The token polling transfer mechanism is as follows: Set M0 to represent the leading large model, M1, M2, ..., M n represent subordinate large models, T represents the speech transfer token (Token), R i represents the response status of the subordinate large model M i , where R i = 1 indicates that the review request can be processed, where R i = 0 indicates that the review request cannot be processed, C i represents the speech content and score of the subordinate large model M i , P(M i ) represents the processing priority of the subordinate large model M i , and k is the maximum number of polling times; At the initial token assignment, let T = M0, that is, the leading large model M0 holds the token. When the leading large model M0 finishes speaking for the first time and the token is passed to the subordinate large models, a random selection scheme is adopted to randomly select the M with the highest priority from the subordinate large models that have not spoken yet. j ; M j = argmax{P(M i )|M i ∈ {M1,...,M n} and M i has not spoken} If R j equals 1, it indicates that at this time M j can process the request, then pass the token T = M j ; If R j equals 0, it indicates that at this time M j cannot process the request, then increment the polling count by 1 and continue to randomly select other silent subordinate large models; If the polling count is greater than the maximum polling count k, then remove M j from the large model expert group; When all subordinate large models have spoken or there are no subordinate large models, the leading large model summarizes its own scores C0, C1, C2,..., C with all subordinate large models n , and calculates the final score, where g is the comprehensive scoring function: C = g(C0, C1, C2, ..., C n )。 3. The method according to claim 1 or 2, characterized in that: After selecting specific business code change records, start to perform the automated code review and evaluation analysis of the multi-expert large model cooperation under the Domain-Driven Design (DDD) architecture.
4. The method according to claim 3, characterized in that: The specific steps are as follows: S10: Detection of business framework integrity Execute a customized code structure detection tool suitable for the Domain-Driven Design (DDD) architecture of the tobacco industry to detect the framework integrity of the changed files, and store the results in the framework non-integrity detection records. S20: Resolution of code dependency relationships Execute a code dependency parsing tool related to the context of the changed code to perform hierarchical dependency parsing, capture the upper and lower layer business codes related to the changed code, and store the relevant business dependencies of the changed code in the business code dependency records. S30: Construction of a multi-expert large model cooperation framework Adopt a multi-expert large model cooperation framework to construct an expert large model code review group, confirm the leading large model and subordinate large models, issue the token controlling the large model's speech to the leading large model, and formulate review rules and prompt templates for the large models, including role requirements, optimization requirements, and specific scoring rules. S40: Final evaluation and warning of the code review system After receiving the changed submitted code, the code review system summarizes the code submission rating, that is, the discrimination is excellent, qualified, and unqualified. The submission records with an unqualified code submission rating will issue a system warning to urge a second code review, facilitating timely code rectification and giving priority to eliminating high-risk vulnerabilities.
5. The method according to claim 4, characterized in that: Code review and evaluation of the leading large model The leading large model conducts code review evaluation according to the framework non-integrity detection record, business code dependency record, changed code details, commit message, and project preference tool class tendency, and gives a code score according to the scoring structure based on business and code standardization requirements; After the review, the code review evaluation results and optimization suggestions are sent to the broadcast common area of the multi-model collaboration framework, and the token is passed to the subordinate large model through the token polling algorithm.
6. The method according to claim 5, wherein Speech token polling algorithm The expert token polling algorithm will detect the subordinate large model currently in the large model code review expert group, randomly select a subordinate large model that has not spoken yet to check the online status of the current subordinate large model, and send a request to check whether the basic requirements for code review can be replied in time. If the subordinate large model returns a response status code indicating that it can reply to the code review request in time, the speech token will be officially passed to the subordinate large model, otherwise the speech token will be passed to the next subordinate large model; If one of the subordinate large models fails to respond correctly multiple times, it will be temporarily excluded from the target subordinate large models of the current round of speech token polling algorithm.
7. The method according to claim 6, wherein Subordinate large model code review evaluation The subordinate large model obtains the relevant information mentioned above through the broadcast common area, analyzes potential problems of the provided changed code according to the prompt words and provides optimization suggestions, and gives a code score under the specified scoring framework; after sending the answer result to the broadcast common area of the multi-model collaboration framework, the token is passed to the next subordinate large model through the token polling algorithm; and so on, until the last subordinate large model conducts code review evaluation and scoring, and the token is passed to the leading large model.
8. The method according to claim 7, wherein Leading large model summary and final score When the last subordinate large model completes the review, the final result is summarized by the leading large model, and the problems and optimization suggestions in the submitted code of this change are summarized item by item. According to the scoring results of all expert large models, the final score result is summarized according to the expert scoring algorithm, and the submitted final score result and evaluation suggestions are sent to the code review system.
Citation Information
Cited By
Intelligent code change recognition system, method and device based on large language model and medium
CN120596354A
Industry report quality evaluation method, apparatus and device, and storage medium
CN121169213A
Large code model training method and electronic equipment
CN121187569A
Illusion mitigation method and device for multi-modal large language model
CN121661442A
A hallucination mitigation method and device of a multi-modal large language model
CN121661442B