Code generation method, equipment, device, medium and product
By obtaining the semantic characteristics of user input information and using emotion analysis model to determine emotional tendency categories, and generating target codes that match users' emotions, the problem of low compatibility between code generation and user needs in existing tools is solved, and the code quality and practicality are improved.
Patent Information
- Application Number
- CN202510709362.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-26
AI Technical Summary
Existing code generation tools are difficult to dig deep into the emotional tendencies and potential needs behind user input information, resulting in a low degree of compatibility with user needs.
By obtaining the semantic characteristics of user input information, using the sentiment analysis model to determine the emotional tendency category, and generating target codes based on the semantic characteristics and emotional tendency category to ensure that the code style matches the user's emotions.
It improves the quality and practicality of code generation, and improves the degree of fit between the target code and user needs by accurately identifying user emotional state and real needs.
Smart Images

Figure CN120540642A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a code generation method, device, apparatus, medium, and product. Background Art
[0002] With the rapid development of information technology, the demand for software development has exploded. Current code generation tools require extremely high levels of developer expertise and experience, significantly limiting development efficiency and code quality. Despite significant progress in automation, code generation tools still struggle to meet user needs. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present application provides a code generation method, device, apparatus, medium and product.
[0004] According to a first aspect of any embodiment of the present application, a code generation method is provided, the method comprising:
[0005] Obtaining user input information and extracting semantic features of the input information;
[0006] Determining, based on the semantic features, a sentiment tendency category that matches the input information;
[0007] Based on the semantic features and the emotional tendency category, a target code is generated, and the coding style of the target code matches the emotional tendency category.
[0008] According to a second aspect of any embodiment of the present application, a code generation device is provided, the device comprising:
[0009] A data processing module is used to obtain user input information and extract semantic features of the input information;
[0010] A sentiment analysis module, configured to determine a sentiment tendency category matching the input information based on the semantic features;
[0011] The code generation module is used to generate target code based on the semantic features and the emotional tendency category, and the code style of the target code matches the emotional tendency category.
[0012] According to a third aspect of any embodiment of the present application, an electronic device is provided, including:
[0013] processor;
[0014] a memory for storing processor-executable instructions;
[0015] The processor implements the method described in any embodiment of the present application by running the executable instructions.
[0016] According to a fourth aspect of any embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the method described in any embodiment of the present application is implemented.
[0017] According to a fifth aspect of any embodiment of the present application, a computer program product is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the method described in any embodiment of the present application is implemented.
[0018] The technical solution provided by this application may have the following beneficial effects:
[0019] According to the above embodiments, by obtaining the user's input information and extracting the semantic features of the input information, the emotional tendency category of the input information is determined according to the semantic features, and the target code is generated based on the semantic features and the emotional tendency category. By performing emotional analysis on the user's input information, the user's emotional state and real needs can be accurately identified, and the semantic features and the user's emotional tendency category can be integrated to ensure the logical accuracy of the target code, thereby improving the degree of fit between the target code and user needs, thereby improving the quality and practicality of the target code.
[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0022] Figure 1 is a flowchart of a code generation method according to an exemplary embodiment of the present application;
[0023] Figure 2 This is a flowchart of a method for determining a target code style according to an exemplary embodiment of the present application;
[0024] Figure 3 This is a flowchart of a method for determining a preference probability value according to an exemplary embodiment of the present application;
[0025] Figure 4 is a flowchart of a method for processing candidate codes according to an exemplary embodiment of the present application;
[0026] Figure 5is a flowchart of another code generation method according to an exemplary embodiment of the present application;
[0027] Figure 6 is a flowchart of a method for generating a universal code element according to an exemplary embodiment of the present application;
[0028] Figure 7 is a flowchart of another code generation method according to an exemplary embodiment of the present application;
[0029] Figure 8 is a structural diagram of an electronic device according to an exemplary embodiment of the present application;
[0030] Figure 9 It is a block diagram of a code generation device according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0032] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0034] Current code development tools typically only provide a superficial understanding of natural language input from users, failing to delve deeper into the underlying emotional tendencies and potential needs. For example, when a user enters "Write a function to calculate the first n terms of the Fibonacci sequence, but I don't know how to use recursion," code development tools often only generate recursive or non-recursive code based on the literal meaning. However, they fail to understand the user's anxiety or distress caused by not knowing how to use recursion, resulting in the generated code being less aligned with the user's needs.
[0035] In order to solve the above problems, this application proposes a code generation method. To further illustrate this application, the following embodiments are provided:
[0036] See also Figure 1 , Figure 1 This is a flowchart illustrating a code generation method according to an exemplary embodiment of the present application. This code generation method can be applied to a code generation system, which can be applied to electronic devices such as mobile phones, computers, digital broadcast terminals, messaging devices, tablet devices, and personal digital assistants. It can also be applied to server-side services such as single servers, cluster servers, and cloud servers. This method can also be executed by other systems or devices in different application scenarios.
[0037] For example, the code generation system can be integrated into code development tools such as an integrated development environment (IDE), an online code editor, an automated testing platform, a code review tool, and a programming education platform; for another example, the code generation system can be integrated into question-and-answer dialogue systems such as intelligent assistants and knowledge base retrieval systems; for another example, the code generation system can be integrated into other systems or devices such as software development collaboration platforms and cloud development environments, and the embodiments of the present application do not limit this.
[0038] like Figure 1 As shown, the code generation method may include the following steps:
[0039] Step 101: Obtain user input information and extract semantic features of the input information.
[0040] In this step, the code generation system may receive input information provided by the user through an input method such as a graphical interface or a command line tool.
[0041] The input information may be text information, and the system may utilize natural language processing technology to perform preprocessing operations on the input information in order to more accurately extract the semantic features of the input information.
[0042] Preprocessing can include tokenization, part-of-speech tagging, and stop word removal. For example, the input text can first be tokenized, breaking down continuous character sequences into meaningful units, such as words or phrases. This step is crucial for subsequent processing and directly impacts the feature extraction model's understanding of the input information. For example, the input information "Write a function to calculate the Fibonacci sequence, but I don't know how to do recursion" can be tokenized into "Write / a / function / to calculate / the / Fibonacci sequence / , / but / I / don't know how to do recursion."
[0043] The system can tag each word with a part of speech, determining its grammatical role in the sentence, such as noun, verb, or adjective. This process helps understand the relationship between words and lays the foundation for further syntactic analysis. For example, "calculate" is tagged as a verb, "Fibonacci sequence" is tagged as a noun, and "will not" is tagged as an adverb.
[0044] The system can remove common words that do not carry important information, such as "of", "is", "one", "I", etc., to reduce noise and improve model efficiency, and obtain "write / function / calculate / Fibonacci sequence / will not / recurse".
[0045] The system uses pre-trained feature extraction models to encode pre-processed input information and extract rich contextual information. Feature extraction models, pre-trained on large-scale datasets, have learned a wealth of linguistic knowledge and patterns, effectively capturing semantic vectors or structured representations in text and inputting semantic features. Examples include the Transformer-based Bidirectional Encoder Representations from Transformers (BERT) model and the Text-to-Text Transfer Transformer (T5) model.
[0046] Feature extraction models can incorporate an attention mechanism to focus on key information points when generating semantic features, such as keywords like "can't" and "recursion." This further improves model performance, not only increasing accuracy but also enhancing system interpretability. Feature extraction models generate one or more semantic features that accurately reflect the intent and needs of user input.
[0047] For example, for the input information "Write a function to calculate the Fibonacci sequence, but I don't know recursion", after the above processing, the semantic features extracted by the system can be expressed as: {"action":"generate","target":"function","task":"calculate Fibonacci sequence"}.
[0048] Among them, action indicates the type of operation the user wants to perform, generate indicates that the user wants to generate new code, target indicates the target code unit the user wants to operate on, function indicates that the user wants to generate a function, task indicates the specific task or functional requirement, that is, the function the user wants the generated code to implement; calculate Fibonacci sequence indicates that the user wants the generated code to be able to calculate the Fibonacci sequence.
[0049] The input information can also be audio information or image information. The system can convert it into text information, and then perform preprocessing operations and semantic feature extraction to provide a basis for subsequent sentiment analysis and code generation.
[0050] In one embodiment, the input information may include current input information and historical input information. Feature extraction may be performed on the current input information to generate a current semantic feature corresponding to the current input information. Historical semantic features corresponding to the historical input information may be obtained. The historical input information may be used as context information for the current input information, and the historical semantic features may be superimposed on the current semantic features to obtain a semantic feature.
[0051] The current input information is the user's current demand description, which can reflect the user's immediate intention and operation goal. The current semantic feature is the semantic feature extracted from the current input information, which can accurately reflect the semantic content of the current input information.
[0052] Historical input information is the input information of the user in the past interaction with the system. It can provide the user's historical background and context information, and help the system understand the relationship between the current input information and previous conversations or operations.
[0053] Historical semantic features are semantic features obtained by extracting features from historical input information. They can reflect the key information and contextual content in historical input information, enabling the system to review and utilize past interaction information, thereby better grasping users' long-term demand patterns and preferences.
[0054] For example, the current input information can be preprocessed to extract key information from the text. The processed text is encoded using a pre-trained feature extraction model to generate a feature vector or structured representation reflecting the semantic content of the current input, namely, the current semantic feature.
[0055] The historical semantic features corresponding to the historical input information can be obtained. The historical semantic features are the semantic features obtained by the system after feature extraction of the historical input information. They can reflect the key information and contextual content in the historical conversation or historical input.
[0056] Historical input information can be used as the context information of the current input information. The historical semantic features can be superimposed on the current semantic features through simple vector splicing, weighted summation, or through special fusion algorithms or neural network layers. Finally, semantic features that comprehensively consider the current input and historical context are obtained, providing richer and more accurate semantic information for subsequent steps such as sentiment analysis and code generation, so that the system can better understand and respond to user needs.
[0057] For example, the current input is "This function doesn't seem right. The result of inputting 4 and 6 is 2, but it should be 1." The system extracts the current semantic feature {"error":"incorrect result for input(4,6)"} from this input. Error indicates an error or problem in the current code, and "incorrect result for input(4,6)" indicates that the user indicated that the function's result was incorrect when the inputs were 4 and 6.
[0058] You can obtain the historical semantic features corresponding to the historical input information "Write a function to calculate the greatest common divisor of two numbers". The historical semantic features are {"action":"generate","target":"function","task":"calculateGCD of two numbers"}. Among them, "calculate GCD of two numbers" indicates that the user wants to generate a function to calculate the greatest common divisor of two numbers.
[0059] Superimpose the historical semantic features with the current semantic features to obtain the final semantic features: {"action":"generate","target":"function","task":"calculate GCD of two numbers","error":"incorrect result for input(4,6)"}.
[0060] As described above, by extracting features from the current input information, generating current semantic features corresponding to the current input information, obtaining historical semantic features corresponding to historical input information, using historical input information as context information for the current input information, and superimposing historical semantic features with current semantic features to obtain semantic features, the comprehensive semantic features provide richer and more accurate information input for code generation, which helps to generate code that better meets user expectations, improves the consistency and accuracy of code generation, reduces repeated user input, and improves the overall quality and efficiency of code generation.
[0061] Step 102: Determine the emotional tendency category that matches the input information based on the semantic features.
[0062] In this step, the code generation system uses a pre-trained sentiment analysis model to perform sentiment classification on the extracted semantic features. Based on deep learning technology, the sentiment analysis model automatically identifies subjective information in text and assigns corresponding sentiment categories. These sentiment categories represent the type of emotion conveyed by the input information and are used to determine the style of the generated code to meet the user's emotional needs.
[0063] Continuing with the input information "Write a function to calculate the Fibonacci sequence, but I don't know how to do recursion" as an example, after preprocessing the input information, the system uses the feature extraction model to encode the preprocessed input information, extract semantic features, and use the attention mechanism to capture sentiment words in the text, such as "can't", so that the feature extraction model pays more attention to this word that expresses the user's confusion.
[0064] The system can input the extracted semantic features into a sentiment analysis model, which, through a softmax classification layer, outputs a sentiment category that matches the input information. For example, based on the semantic features of the sentence "Write a function to calculate the Fibonacci sequence, but I don't know how to use recursion," the sentiment analysis model outputs a negative sentiment category, indicating that the user is highly confused about recursion.
[0065] Step 103: Generate target code based on the semantic features and the sentiment tendency category, where the coding style of the target code matches the sentiment tendency category.
[0066] In this step, the system can utilize code generation models and pre-defined code library retrieval methods to generate the corresponding code structure and logic based on key information from the semantic features. Based on the sentiment category, the generated code structure and logic can be adjusted to produce the target code. The generated target code not only implements the user's desired functionality, but also matches the sentiment category with its coding style, thereby enhancing the code's personalization.
[0067] The target code is the code that meets the user's needs and emotional tendencies. The code generation model can be trained based on Generative Adversarial Networks (GAN), Sequence to Sequence (Seq2Seq) models, or other deep learning architectures.
[0068] Continuing with the example of the input message "Write a function to calculate the Fibonacci sequence, but I don't know recursion," the system can input semantic features and sentiment categories (negative sentiment) into the code generation model. The code generation model generates the code structure and logic of the corresponding Fibonacci sequence function based on the key information in the semantic features.
[0069] The code generation model can generate target code with detailed comments and example information based on negative sentiment to help users understand and use it. The generated target code can be as follows (using Python as an example):
[0070]
[0071]
[0072] Through this code generation process, the system provides users with code generation solutions that better suit their emotional state. For example, it provides a non-recursive Fibonacci sequence calculation method, accompanied by detailed comments and examples to help users understand and solve problems. The generated code not only implements the user's desired functionality but also helps users overcome confusion about recursion through detailed comments and examples, thereby improving code readability and ease of use.
[0073] The code generation method of this embodiment obtains the user's input information and extracts the semantic features of the input information. According to the semantic features, the emotional tendency category of the input information is determined. Based on the semantic features and the emotional tendency category, the target code is generated. By performing emotional analysis on the user's input information, the user's emotional state and real needs can be accurately identified. By integrating the semantic features and the user's emotional tendency category, the degree of fit between the target code and the user's needs is improved on the basis of ensuring the logical accuracy of the target code, thereby improving the quality and practicality of the target code.
[0074] In the aforementioned embodiments, we described how to receive user input, extract its semantic features, and use a sentiment analysis model to determine the sentiment category. Ultimately, based on the semantic features and sentiment category, we generated target code that conforms to a specific coding style. The following embodiments provide a more detailed description of the target code generation process, which is applicable to any of the above embodiments.
[0075] In one embodiment, when generating a target code, first, the target code style of the target code can be determined based on the emotional tendency category; next, the initial code is generated based on the semantic features; finally, the code expression form of the initial code is adjusted according to the target code style to generate the target code.
[0076] The target code style is the presentation format and characteristics of the target code, ensuring that the generated target code meets user emotional needs, such as the level of detail in comments and the complexity of the code structure. The initial code is a preliminary code generated based on semantic features and can serve as the basis for generating the target code, which can be subsequently adjusted based on the target code style.
[0077] The code representation is the way the target code is presented. This can include at least one of the following: comments and logical structure. Logical structure is the organization and flow of the various parts of the code, such as function definitions, implementation methods, and calling order.
[0078] For example, the target code style can be determined based on the sentiment categories that match the input information by searching a preset style library or deep learning models. Based on semantic features, a logical framework for the code can be constructed, translating the user's natural language requirements into structured, executable initial code. The preset style library pre-stores mappings between different sentiment categories and different code styles.
[0079] Adjust the initial code's comments, logical structure, and other aspects based on the target code style. For example, if the target code style is detailed comments, add detailed comments and examples to the initial code to ensure that users can clearly understand the functionality and logic of each part of the code. You can also optimize the code's logical structure to make it easier to read and maintain.
[0080] To further introduce the process of determining the target coding style, Figure 2 A flow chart of a method for determining a target code style is shown. The method may include the following steps:
[0081] Step 201: Receive input information input by the user.
[0082] In this step, the user input information may be received through the user interaction interface, for example, the input information is "this tool is very useful".
[0083] Step 202: Perform preprocessing and feature extraction on the input information, such as word segmentation, part-of-speech tagging, and stop word removal, to obtain semantic features.
[0084] In this step, the input information can be preprocessed, including operations such as word segmentation, part-of-speech tagging, and stop word removal. The feature extraction model encodes the preprocessed input information and extracts semantic features. For example, if the input information "This tool is very useful" is input into the feature extraction model, the feature extraction model will output a 768-dimensional semantic feature.
[0085] During feature extraction, the feature extraction model can use the attention mechanism to capture key sentiment words in the input information. For example, the feature extraction model pays more attention to the keyword "easy to use".
[0086] Step 203: Use the sentiment analysis model to perform sentiment classification on the semantic features to determine the sentiment tendency category that matches the input information.
[0087] In this step, the sentiment analysis model is used to perform sentiment classification on the semantic features. The sentiment analysis model outputs positive, negative, neutral and other sentiment tendency categories that match the input information through the Softmax classification layer.
[0088] Step 204: Determine the target code style of the target code according to the sentiment tendency category.
[0089] In this step, a code style corresponding to the emotional tendency category can be retrieved from a preset style library as a target code style of the target code.
[0090] The system can continuously monitor user input and parse it in real time. For example, if a user responds to a target code with "the code is too difficult to understand," the system uses natural language processing to identify the user's dissatisfaction with the target code's readability. This triggers a code style optimization mechanism, adjusts the target code style, generates annotated code, or changes the code's logical structure. Furthermore, through reinforcement learning, the system continuously optimizes the code style and improves code generation.
[0091] As described above, by determining the target code style according to the emotional tendency category, generating the initial code based on the semantic features, and then adjusting the code expression of the initial code, the generated target code not only meets the functional requirements but also conforms to the user's personalized style preferences, thereby improving the code quality and user experience.
[0092] In one embodiment, when determining the emotional tendency category that matches the input information, a probability value may be assigned to each emotional tendency category based on semantic features.
[0093] When determining the target code style of the target code, the emotional probability values of different code styles can be determined based on the probability values of each emotional tendency category; secondly, the user's preference probability values for different code styles can be determined based on the user's historical behavior data; then, the emotional probability value and the preference probability value are fused according to the preset weights to generate the final probability values of different code styles; finally, the code style with the highest final probability value is used as the target code style.
[0094] The sentiment probability value is the probability of the coding style corresponding to each sentiment category, reflecting the likelihood that the input information belongs to a specific coding style. The preference probability value quantifies the user's historical preference for different coding styles. The final probability value is a composite probability value that combines the sentiment probability value and the preference probability value and is used to select the final target coding style.
[0095] Historical behavior data refers to the operation records and behavior information left by users when using code generation systems or other related development tools. It is used to analyze users' programming habits, technical preferences, and ability levels. For example, it includes past user input, feedback on generated code, modification habits, commonly used programming languages, preferred coding styles, frequently used library functions, and types of code generation tasks.
[0096] For example, the system can use the sentiment analysis model to classify the sentiment tendencies of the input information and assign a probability value to each sentiment tendency category. The probability value can be between 0 and 1. The higher the probability value, the greater the possibility that the input information matches the sentiment tendency category, such as the probability value of negative sentiment is 0.9, the probability value of neutral sentiment is 0.05, and the probability value of positive sentiment is 0.05.
[0097] Based on the mapping between sentiment categories and coding styles, the sentiment probability of each coding style is determined based on the probability of each sentiment category. If negative sentiment corresponds to a detailed comment style, neutral sentiment corresponds to a standard programming style, and positive sentiment corresponds to a concise style, then the sentiment probabilities of the corresponding coding styles are 0.9, 0.05, and 0.05, respectively.
[0098] The system first collects various behavioral data from users throughout their past use, such as whether they frequently use Python or Java as their programming language, whether they prefer a concise or detailed coding style, and which specific libraries, such as Flask or NumPy, they frequently use in their code. The system also records users' adoption of previously generated target code and subsequent modifications.
[0099] By deeply analyzing this historical behavior data, the system can extract characteristics of a user's coding style, such as the density and complexity of code comments, and count the frequency of users' use of specific library functions, thereby understanding their development habits. Furthermore, the system can identify patterns in user code modifications, such as whether users tend to delete redundant comments.
[0100] Based on this analysis, the system uses machine learning algorithms to build a user behavior model, which then predicts the probability of a user's preference for different coding styles. This model categorizes user preferences into various coding styles, such as concise code, well-commented code, and efficient code. For each style, it calculates a probability value between 0 and 1 to quantify the user's preference for that style.
[0101] If the user's preference probability value for the detailed comment style is 0.7, the preference probability value for the standard programming style is 0.2, and the preference probability value for the concise style is 0.1.
[0102] The sentiment and preference probabilities are combined according to preset weights to generate the final probabilities for different coding styles. The preset weights can be set to 70% for the sentiment probability and 30% for the preference probability, or 50% each for the sentiment and preference probability, and so on.
[0103] For example, assuming a weight of 70% for sentiment probability and 30% for preference probability, the final probability of the fused detailed comment style is 0.9 × 0.7 + 0.7 × 0.3 = 0.84, the final probability of the standard programming style is 0.05 × 0.7 + 0.2 × 0.3 = 0.095, and the final probability of the concise style is 0.05 × 0.7 + 0.1 × 0.3 = 0.06. The coding style with the highest final probability is the detailed comment style, which can be used as the target coding style.
[0104] To further introduce the process of determining the preference probability value, Figure 3 A flow chart of a method for determining a preference probability value is shown. The method may include the following steps:
[0105] Step 301: Collect user's historical behavior data.
[0106] In this step, the system collects historical behavior data of users in code development tools. The historical behavior data includes but is not limited to: commonly used programming languages (such as Python, Java), commonly used code styles (such as concise code style, detailed comment style), preferred library functions (such as Flask, NumPy), and historical generated code adoption and modification records.
[0107] Step 302: Perform static analysis on the historical behavior data to obtain static analysis results.
[0108] In this step, the collected historical behavior data can be statically analyzed to analyze the format, style, and grammatical structure of the user code, and to extract information such as code style features (such as comment density and code complexity) and library function usage frequency.
[0109] Step 303: Determine whether a code pattern is identified in the static analysis results.
[0110] In this step, you can determine whether there are identifiable code patterns in the static analysis results, such as common design patterns, algorithm implementations, etc.
[0111] If a code pattern is identified in the static analysis results, then proceed to step 304;
[0112] If no code pattern is identified in the static analysis results, step 305 is executed.
[0113] Step 304: Query the practice library.
[0114] In this step, if a code pattern is identified, the system uses a pre-trained large model to search the practice library for vectors matching the pattern and generates user behavior features based on the matching vectors. The practice library is a collection of proven, efficient, and high-quality code snippets, solutions, and design patterns. It can be written and maintained by domain experts or experienced developers, providing high-quality best practices.
[0115] Step 305: Perform syntax check on the historical code data.
[0116] In this step, if the code pattern is not identified, the system can use the abstract syntax tree and preset rules to perform syntax detection on the historical code data and generate user behavior features based on the detection results.
[0117] Step 306: Establish a user behavior model.
[0118] In this step, the static analysis results and the generated user behavior features can be input into a machine learning algorithm (such as clustering and classification) to build a user behavior model that can predict the user's programming preferences.
[0119] Step 307: Using the user behavior model, determine the probability values of the user's preference for different coding styles.
[0120] In this step, the probability values of the user's preference for different coding styles can be determined based on the user behavior model, thereby quantifying the user's preference for the coding style.
[0121] As mentioned above, by assigning probability values to each sentiment tendency category based on semantic features, determining the sentiment probability values of different code styles, and determining the preference probability values of different code styles based on the user's historical behavior data, and integrating the sentiment probability values with the preference probability values, we can dynamically determine the target code style, accurately match the user's style preferences, and improve the personalization level of code generation and user satisfaction.
[0122] In one embodiment, when the emotional tendency category is negative emotional tendency, the negative coding style is used as the target coding style; when the emotional tendency category is neutral emotional tendency, the neutral coding style is used as the target coding style; when the emotional tendency category is positive emotional tendency, the positive coding style is used as the target coding style.
[0123] The information richness of auxiliary information for negative, neutral, and positive coding styles decreases. Auxiliary information is information that helps users understand and use the code, improving its readability and usability. Examples include comments, function call examples, and parameter descriptions.
[0124] The negative code style has the most auxiliary information, which can include detailed code comments, example calls, detailed support documents, simplified code logic, etc. The negative code style can be a detailed comment style; the neutral code style has relatively moderate auxiliary information, providing basic comments and instructions. The neutral code style can be a standard programming style, providing basic comments, and following best practices; the positive code style has the least auxiliary information, mainly focusing on the simplicity and efficiency of the code, reducing redundant comments, and the comments are relatively concise and clear. The positive code style can provide performance optimization suggestions. The positive code style can be a concise programming style.
[0125] Negative sentiment reflects the emotional tendency of input information that conveys negative emotions, such as difficulty, incompetence, disappointment, error, poor quality, or repetitive context. Neutral sentiment reflects the emotional tendency of input information that lacks emotional overtones. Input information that lacks emotional overtones can include factual descriptions or instructions. Positive sentiment reflects the emotional tendency of user input information that conveys positive emotions, such as good, excellent, efficient, satisfactory, and enjoyable.
[0126] As mentioned above, by directly mapping different emotional tendency categories to specific code styles, the style determination process can be simplified, ensuring that the generated code style is consistent with the user's emotional state, and improving the readability and practicality of the code.
[0127] In one embodiment, the initial code can be determined by searching a preset code library. A set of candidate codes matching the semantic features is obtained from the preset code library. The initial code can be adjusted by obtaining candidate codes matching the target code style from the candidate code set and adjusting the candidate codes based on the target code style.
[0128] The candidate code set is a list of multiple candidate codes that match the semantic features, retrieved from a preset code library, such as multiple versions of a function that implements the Fibonacci sequence. A candidate code is a single code option in the candidate code set, such as a Fibonacci sequence function that uses iteration.
[0129] For example, the system can obtain a set of candidate codes that match the semantic features from a preset code library. For example, a similarity algorithm such as cosine similarity can be used to determine one or more candidate codes that match the semantic features, and the one or more candidate codes can be combined into a candidate code set.
[0130] The system can also combine grammatical rules, best practices and other rule engines to screen candidate codes in the candidate code set and filter out invalid or low-quality candidate codes. The system can use the candidate codes in the candidate code set as the initial code.
[0131] Exemplarily, the system can obtain a set of candidate codes that match semantic features in a preset code library. For example, the user inputs "implement a function to calculate the greatest common divisor of two numbers". The system can extract the semantic features corresponding to the input information, and combine the project environment (such as Python language, imported math library) and historical data (such as users' common use of math.gcd function, preference for concise code) to find matching code snippets in the code library, determine multiple candidate codes through the cosine similarity algorithm, and form a candidate code set. The candidate codes may include:
[0132] The first implementation uses math.gcd:
[0133] import math
[0134] def calculate_gcd(a,b):
[0135] return math.gcd(a,b)
[0136] The second implementation method uses the Euclidean algorithm:
[0137]
[0138] The system uses a rule engine that combines grammatical rules and best practices to filter out invalid or low-quality candidate code. Based on the target coding style (for example, if a user frequently uses math.gcd and prefers a concise coding style), the system ultimately selects candidate code from the candidate code set that matches the target coding style and prioritizes the concise implementation of math.gcd.
[0139] Candidate codes that match the clean coding style:
[0140]
[0141] The candidate code can be adjusted based on the target code style, for example, by appropriately deleting redundant comments and further simplifying the code's logical structure, to obtain the target code. This code generation method ensures that the generated initial code not only meets functional requirements but also conforms to the user's coding style preferences, effectively improving the code's usability and user experience.
[0142] If no candidate code matching the semantic features exists in the pre-set code base, the code generation model can be used to generate initial code based on the semantic features. The system can store the generated initial code in the pre-set code base, allowing for rapid retrieval and utilization when similar requirements are encountered later, improving code generation efficiency while continuously enriching and improving the code base resources.
[0143] To further introduce the process of determining candidate codes, Figure 4 A flowchart of a method for processing candidate codes is shown. The method may include the following steps:
[0144] Step 401: Retrieve code data matching the semantic feature from a preset code library.
[0145] In this step, the system can search for matching code snippets or library functions and other code data in a preset code library based on the extracted semantic features.
[0146] Step 402: Determine whether matching code data is retrieved.
[0147] In this step, the system determines whether matching code data is retrieved from the preset code library.
[0148] If matching code data is retrieved, continue to step 403;
[0149] If no matching code data is found, step 405 is executed.
[0150] Step 403: Use a similarity algorithm to sort the multiple code data and generate a candidate code set.
[0151] In this step, a similarity algorithm may be used to sort the retrieved multiple code data, and a candidate code set may be generated based on the similarity scores.
[0152] Step 404: Evaluate and screen the candidate codes in the candidate code set to determine an initial code.
[0153] In this step, the candidate codes in the candidate code set can be evaluated and screened in combination with preset rules to filter out invalid or low-quality candidate codes, and the candidate codes in the screened candidate code set are used as initial codes.
[0154] Step 405: Generate initial code based on semantic features using the code generation model.
[0155] In this step, if there is no code data matching the semantic features in the preset code library, the code generation model can be used to generate the initial code based on the extracted semantic features. When generating the initial code, the code generation model can perform function call and definition checks on the initial code. When generating function calls, it checks whether the corresponding definitions exist. The initial code can also be checked for variable declarations. When generating variable usage, it checks whether the variable has been declared. For example, when generating result = calculate_fibonacci (n), it checks whether result has been declared. It can also check the dependency library: when generating code, it checks whether the required dependency library has been imported. For example, when generating flask.Flask(), it checks whether the flask library has been imported.
[0156] After the initial code is generated, it can be stored in a preset code library so that it can be quickly retrieved and utilized when similar requirements are encountered in the future.
[0157] As described above, by searching a preset code library and selecting candidate codes that match the target code style for adjustment, high-quality initial code can be generated efficiently, saving code generation time and computing resources while maintaining consistency in code style.
[0158] In one embodiment, after the initial code is generated, a syntax tree detection technology may be used to perform syntax detection on the initial code. In response to detecting that the initial code has syntax problems, the initial code is corrected to obtain a corrected initial code.
[0159] Exemplarily, the code generation model may include a syntax checking unit. After the code generation model generates initial code based on the extracted semantic features, the syntax checking unit may be used to perform syntax checking on the initial code by constructing an Abstract Syntax Tree (AST). The AST is a tree-structured representation of the source code that can reflect the grammatical structure of the initial code. The syntax checking unit, based on a discriminator constrained by the AST syntax tree, can identify syntax errors in the initial code, such as incorrect syntax, incomplete structure, etc.
[0160] In response to detecting that there is a syntax problem in the initial code, the code generation model can accurately locate the error location based on the structural information of the AST, and automatically correct it to generate a corrected initial code that conforms to the syntax specification.
[0161] As mentioned above, through syntax tree detection technology and initial code correction, syntax errors can be automatically detected and corrected, the legality and compilability of code generation can be improved, and the user's debugging costs can be reduced.
[0162] To further introduce the target code generation process, Figure 5 A flow chart of another code generation method is shown. The code generation method may include the following steps:
[0163] Step 501: Extracting semantic features of input information.
[0164] In this step, the system can preprocess and extract features of the user's input information to extract the semantic features of the input information.
[0165] Step 502: Assign a probability value to each sentiment tendency category according to the semantic features.
[0166] In this step, the sentiment analysis model can be used to analyze the extracted semantic features and assign a probability value to each preset sentiment tendency category (such as negative sentiment tendency, neutral sentiment tendency, positive sentiment tendency), indicating the possibility that the input information belongs to the sentiment tendency category.
[0167] Step 503: Determine the emotion probability values of different coding styles based on the probability values of each emotion tendency category.
[0168] In this step, based on the preset mapping relationship between sentiment tendency categories and coding styles, for example, the coding style corresponding to negative sentiment tendency is a negative coding style, the coding style corresponding to neutral sentiment tendency is a neutral coding style, and the coding style corresponding to positive sentiment tendency is a positive coding style. The probability value of each sentiment tendency category is converted into the sentiment probability value of the corresponding coding style.
[0169] Step 504: Determine the user's preference probability values for different coding styles based on the user's historical behavior data.
[0170] In this step, we collect and analyze historical user behavior data, such as commonly used programming languages, coding styles, and library function preferences, to build a user behavior model. Using this model, we can predict user preferences for different coding styles and calculate the probability of preference.
[0171] Step 505: The emotion probability value and the preference probability value are fused according to preset weights to generate final probability values of different coding styles.
[0172] In this step, the emotion probability value obtained in step 503 and the preference probability value obtained in step 504 are fused according to preset weights to calculate the final probability value of each coding style.
[0173] Step 506: The coding style with the highest final probability value is used as the target coding style.
[0174] In this step, the code style with the highest final probability value can be selected as the target code style to ensure that the generated code conforms to the user's emotional tendencies and historical preferences.
[0175] Step 507: Generate an initial code based on the semantic features.
[0176] In this step, the code generation model can be used to generate initial code that meets the functional requirements based on the extracted semantic features. Alternatively, a set of candidate codes that match the semantic features can be retrieved from a pre-set code library, and the candidate codes that match the target code style can be selected as the initial code.
[0177] Step 508: Perform syntax checking on the initial code using syntax tree checking technology.
[0178] In this step, an abstract syntax tree is constructed by a syntax checking unit, and syntax checking is performed on the initial code to identify syntax errors therein.
[0179] Step 509: In response to detecting that the initial code has a grammatical problem, the initial code is corrected to obtain a corrected initial code.
[0180] In this step, if a syntax problem is detected in the initial code, the code generation model can accurately locate the error location based on the structural information of the AST, automatically correct it, and generate the corrected initial code.
[0181] Step 510: According to the target code style, adjust the code representation of the initial code to generate the target code.
[0182] In this step, according to the determined target code style, the code expressions such as comments and logical structure of the revised initial code are adjusted, and finally the target code that conforms to the target code style is generated.
[0183] In the aforementioned embodiments, we described how a sentiment analysis model identifies the emotional tendency of input information, combines historical user behavior data to determine a target code style, then generates initial code based on semantic features, performs syntax checking and corrections on the initial code, and ultimately dynamically adjusts the target code's code representation based on the target code style. The following embodiments provide a more detailed description of the target code adjustment process, which can be applied to any of the above embodiments.
[0184] In one embodiment, the target code can be grammatically converted according to the mapping rules between different programming languages to obtain the converted code of the target programming language. The code structure of the converted code is adjusted based on the programming rules of the target programming language to generate the adapted code.
[0185] The code structure may include at least one of the following: code format, dependency declaration, conversion code that conforms to the syntax requirements of the target programming language, and adaptation code that conforms to the programming rules of the target programming language.
[0186] For example, the system can determine the target code's initial programming language by analyzing historical conversations, examining project configuration files, and other methods. The target programming language can also be determined by receiving user instructions and inferring the user's preferred language. The system integrates multiple language adapters supporting different programming languages, which can be called through a unified interface to perform syntax conversion and structural adjustments on the target code.
[0187] During the syntax conversion process, the grammatical structures in the target code can be converted into equivalent implementations in the target programming language based on predefined mapping rules between different programming languages (such as Python, Java, C++, JavaScript, etc.). For example, when converting from Python to C++, Python list derivations can be converted into C++ Standard Template Library (STL) vector container operations, while ensuring that basic syntax elements such as loop structures and conditional judgments conform to the syntax specifications of the target programming language. For complex structures such as function definitions and class declarations, a complete syntax tree mapping relationship can be established to ensure that the converted code is functionally equivalent to the target code.
[0188] The code structure of the conversion code can be adjusted based on the programming rules of the target programming language. For example, for Python language conversion code, the indentation, naming conventions and line length of the conversion code can be adjusted according to the PEP 8 specification; for Java language conversion code, the Java Code Conventions standard can be used to handle the position of curly braces and package structure; for C++ language conversion code, the memory management strategy and template usage can be optimized, and C++ naming conventions such as camel case naming and semicolon ending can be followed.
[0189] It can also automatically generate the dependency declarations required by the target programming language, such as Python's import statement, Java's package declaration, and C++'s header file inclusion directives. For example, when a user needs to implement Fibonacci sequence calculations, the system not only correctly uses vector containers and for loop structures, but also automatically adds the necessary #include directives, names variables and functions according to C++ naming conventions, and ultimately generates compatible code.
[0190] As mentioned above, by supporting multi-language adaptation of the target code, syntax conversion and structural adjustment are performed according to the rules of different programming languages, the flexibility and applicability of code generation are improved to meet diverse development needs.
[0191] In one embodiment, the target code may be subjected to an anomaly detection to obtain a detection result. If the detection result is an abnormal result, an abnormality prompt message and a correction code for the target code are generated. Based on the correction instruction input by the user, the target code is corrected using the correction code or the code contained in the correction instruction.
[0192] The detection results are used to indicate whether there are exceptions or problems in the target code, such as syntax error detection results, runtime exception detection results, etc. Abnormal results are used to indicate that there are exceptions or problems in the target code, such as a null pointer exception when the code is running.
[0193] The exception prompt information provides a detailed description and suggestions of the abnormal problems in the target code, which is used to help users understand and solve the problems in the code. For example, the variable on line 10 is undefined, and using Redis cache tokens can increase QPS by 3 times.
[0194] Correction codes are code snippets used to fix abnormal problems in the target code. They are used to provide solutions and help users correct the code, for example, adding a code snippet for variable declaration.
[0195] For example, static analysis, dynamic testing, and other methods can be used to detect anomalies in the target code, identify possible syntax errors, logical flaws, or performance issues, and generate detection results that include the location and type of the anomaly. If the detection result is an anomaly, a corresponding anomaly prompt message can be generated, detailing the error type, scope of impact, and recommended repair methods, and directly applicable corrected code can be generated as a reference solution.
[0196] Users can choose to accept the corrected code or enter a custom code based on their needs. The system then updates the target code based on the corrected code or the code contained in the corrected instructions. After the correction is complete, the system automatically performs a secondary verification to ensure that the issue has been effectively resolved. The entire detection and correction process can be executed repeatedly until all anomalies are completely resolved. The system can also output the corrected code version and provide a detailed modification record and optimization suggestions.
[0197] As described above, by performing exception detection and correction on the target code, potential code problems can be discovered and resolved in a timely manner, and detailed exception prompt information and correction suggestions can be provided to users, helping users to quickly optimize the code and thus improve development efficiency.
[0198] In one embodiment, the target code is subjected to anomaly detection, including at least one of the following: detecting syntax problems of the target code; detecting data type matching problems of the target code; detecting code quality problems of the target code based on a preset rule engine; running the target code in a sandbox environment to detect abnormal running data of the target code; and performing unit testing on the target code.
[0199] Data type mismatching is a problem involving inconsistent or incompatible data types within the target code, such as assigning a string value to an integer variable. The pre-defined rule engine is a set of predefined rules used to monitor the quality and potential issues of the target code. It identifies common errors and areas for improvement, such as unused variables and unhandled exceptions.
[0200] Code quality issues are those that exist in the target code that do not conform to coding standards or best practices, such as overly long functions or unclear variable names. Runtime exception data is abnormal data generated during the execution of the target code, indicating errors or abnormal conditions that occurred during code execution, such as array out-of-bounds errors and divide-by-zero exceptions.
[0201] For example, technologies such as abstract syntax trees and symbolic execution can be used to check the structural integrity of the target code and identify basic syntax issues such as bracket matching and indentation format. Type inference can be used to detect data type mismatches in the target code, verifying data type consistency in variable assignments and function calls, for example, checking whether an int variable is incorrectly assigned to a string. Based on a pre-set rule engine, the target code can be checked for code quality issues such as unused variables and unhandled exceptions to ensure compliance with coding standards.
[0202] Sandbox isolation technology can be used to run the target code in a sandbox environment and detect operational anomalies, such as real-time capture of null pointer references, array out-of-bounds errors, and other operational anomalies. An automated unit testing framework can also be used to automatically generate test cases, which can be used to unit test the target code and verify the correctness of the code's functionality. For example, this can be used to simulate HTTP request tests, verify the output accuracy of mathematical calculation functions, or handle boundary conditions in business logic.
[0203] For example, the generated target code could be:
[0204] def divide(a,b):
[0205] return a / b
[0206] result = divide(10,0)
[0207] By running the target code in a sandbox environment, the run-time exception data "Line 3: Divide by zero error" indicating a divide by zero error is detected. The user can be provided with an exception message "Suggest adding divisor check" and the corrected code:
[0208] def divide(a,b):
[0209] if b == 0:
[0210] raise ValueError("divisor cannot be zero")
[0211] return a / b
[0212] result = divide(10,0)
[0213] In response to receiving the correction instruction input by the user, the target code is corrected using the correction code, which can shorten the debugging cycle of the target code and improve development efficiency.
[0214] As mentioned above, through multi-dimensional anomaly detection methods such as syntax problem detection, data type matching detection, code quality detection, runtime anomaly detection and unit testing, the correctness and reliability of the generated target code can be ensured, and the code quality can be improved in all aspects.
[0215] In one embodiment, a universal code element is generated based on the target code's runtime environment data and / or code call requirements of the input information in the runtime environment, and the universal code element is merged with the target code.
[0216] The runtime environment data refers to the environment in which the target code runs. It can include at least one of the following: project environment resources and historical code data. Project environment resources are the resources required for project operation and provide the basic environment for code execution, such as programming languages, frameworks, and dependency libraries. Historical code data refers to existing code resources in the current project, which can help understand the existing structure and functionality of the project and avoid duplication of development. For example, historical code data includes defined functions, classes, and variables.
[0217] A code call request is a user-entered requirement for a specific code function, such as a user's desire to call the defined calculate_fibonacci function. Common code elements are common code snippets generated based on runtime environment data and / or code call requirements. These elements, such as import statements and initialization code, enhance the integrity and usability of the code.
[0218] The system can analyze the project environment resources of the current project where the target code is located, such as programming languages, frameworks, dependent libraries, etc., as well as historical code data, such as existing function definitions, variable declarations, etc. The input information is parsed through natural language processing technology (such as BERT), the user's intention is extracted, and the context is understood in combination with the historical conversation content to determine the code call requirements of the input information in the running environment. For example, if the user's input information is "call a function to calculate the Fibonacci sequence", the determined code call requirement is that the user wants to call the defined calculate_fibonacci function.
[0219] The system searches the pre-set codebase for code snippets or library functions that best match the runtime environment data and / or code call requirements, and performs context matching. For example, if the input is "implement a login interface using the OAuth 2.0 authorization code model," the pre-set codebase will match OAuth 2.0 implementation code related to the Flask framework.
[0220] Dynamic code completion automatically generates necessary declarations, initialization, and dependency code based on the context, identifying common code elements and ensuring logical coherence. For example, it generates import statements for the Flask library or installation instructions for related dependent libraries.
[0221] Ultimately, the system will merge the generated common code elements with the target code (such as the Python code of the Flask framework, including JWT token generation) to obtain complete code that meets the requirements of the project environment, thereby improving the practicality and accuracy of code generation and ensuring that the generated code can be seamlessly integrated into the current project.
[0222] To further introduce the process of generating common code elements, Figure 6 A flow chart of a method for generating a general code element is shown. The method may include the following steps:
[0223] Step 601: Determine whether the target code's operating environment data is obtained.
[0224] In this step, the system can check whether the target code's operating environment data is obtained.
[0225] If the operating environment data is obtained, continue to step 602;
[0226] If the operating environment data is not obtained, step 603 is executed.
[0227] Step 602: Identify common code elements in the knowledge base and past cases.
[0228] In this step, the system can search for common code elements such as code snippets, function templates, class definitions, etc. that match the operating environment data in the knowledge base and past cases of the current project as the basic components to improve the target code.
[0229] Step 603: Analyze the user's intention on the input information and determine the code calling requirement.
[0230] In this step, the system uses natural language processing technology to parse the user's input information, extract the user's intention, and combine historical input information for contextual understanding to determine the code call requirements of the input information.
[0231] Step 604: Generate common code elements according to code calling requirements.
[0232] In this step, corresponding common code elements are generated based on the code call requirements determined in step 603. For example, if the calculate_fibonacci function needs to be called, a code snippet for the function call can be generated, and the snippet can be ensured to be logically coherent with the existing code in the project environment.
[0233] Step 605: Merge the common code elements and the target code.
[0234] In this step, the generated common code elements are integrated with the existing target code. For example, necessary syntax and logic adjustments are made to the target code, dependencies are handled, and the target code and common code elements are inserted into appropriate locations to ensure the integrity and correctness of the final code. For example, the system can ensure that the calculate_fibonacci function is correctly defined and all necessary dependent libraries are imported before calling it. The resulting code not only meets the functional requirements but also integrates seamlessly into the current project.
[0235] As described above, by generating common code elements based on runtime environment data and / or code call requirements and integrating them with the target code, more complete and practical code can be generated, ensuring that the code can be integrated into the current project, thereby improving the usability of the code.
[0236] To further introduce the target code generation process, Figure 7 A flow chart of another code generation method is shown. The code generation method may include the following steps:
[0237] Step 701: Receive user input information.
[0238] In this step, the system may receive input information from the user through the user interface, for example, "implement a function to calculate the greatest common divisor of two numbers."
[0239] Step 702: Extract semantic features of input information.
[0240] In this step, the system performs preprocessing operations such as word segmentation, part-of-speech tagging, and stop word removal on the input information. For example, the input information "Implement a function to calculate the greatest common divisor of two numbers" is segmented into "Implement / a / function / calculate / the greatest common divisor of / two numbers".
[0241] Using the pre-trained feature extraction model, we encode the pre-processed input information and extract semantic features. For example, the generated semantic features can be represented as: {"action":"generate","target":"function","task":"calculate GCD of two numbers"}.
[0242] Step 703: Using the sentiment analysis model, assign a probability value to each sentiment tendency category.
[0243] In this step, the sentiment analysis model classifies the semantic features and assigns a probability value to each sentiment category (e.g., negative, neutral, positive). For example, if the input contains a sentence expressing confusion, such as "But I don't know how to do recursion," the sentiment analysis model will output a negative sentiment with a probability value of 0.9.
[0244] If the input is "Login interface implementing the OAuth 2.0 authorization code model" and contains no emotionally charged words, the sentiment analysis model outputs a neutral sentiment with a probability of 0.8. If the input contains phrases expressing satisfaction or excitement, such as "This tool is very useful!", the sentiment analysis model outputs a positive sentiment with a probability of 0.9.
[0245] Step 704: Determine the emotion probability values of different coding styles based on the probability values of each emotion tendency category.
[0246] In this step, based on the mapping relationship between sentiment categories and coding styles, the probability values of each sentiment category are converted into sentiment probability values for the corresponding coding styles. For example, a negative sentiment category corresponds to a negative coding style, a neutral sentiment category corresponds to a neutral coding style, and a positive sentiment category corresponds to a positive coding style. If the probability value of a negative sentiment category is 0.9, then the sentiment probability value of a negative coding style is 0.9.
[0247] Step 705: Determine preference probability values for different coding styles based on the user's historical behavior data.
[0248] In this step, we collect and analyze historical user behavior data, such as commonly used programming languages, coding styles, and library function preferences, to build a user behavior model. Based on this model, we can predict the user's preference for different coding styles and calculate the probability of preference. For example, a user's historical data indicates a probability of 0.7 for a preference for a positive coding style.
[0249] Step 706: Determine the target coding style by combining the sentiment probability value and the preference probability value.
[0250] In this step, the sentiment probability value and the preference probability value can be combined according to the preset weights to calculate the final probability values of different coding styles. The coding style with the highest final probability value is selected as the target coding style.
[0251] Step 707: Determine the initial code based on the acquired candidate code set.
[0252] In this step, code data that matches the semantic features is retrieved from the preset code library. If multiple code data are retrieved, the code data are sorted using a similarity algorithm to generate a candidate code set. The candidate code in the candidate code set is used as the initial code.
[0253] Step 708: Perform syntax check and correction on the initial code.
[0254] In this step, the initial code is syntax-checked by constructing an abstract syntax tree (AST) to identify potential syntax errors. If a syntax error is detected, the error is pinpointed based on the structure of the AST and corrected, generating a corrected initial code that conforms to the syntax specifications.
[0255] Step 709: According to the target code style, adjust the code representation of the initial code to generate the target code.
[0256] In this step, candidate codes that match the target code style are retrieved from the candidate code set. Based on the determined target code style, the candidate codes' code representation, such as comments and logical structure, are adjusted to generate the target code. For example, if the target code style is a detailed comment style, the system will add detailed comments and example information to the candidate code to ensure that users can clearly understand the function and logic of each part of the code.
[0257] Step 710: Merge the common code elements and the target code.
[0258] In this step, the project's runtime environment data and historical code data are analyzed to generate common code elements. For example, if the project uses Python and the Flask framework, corresponding import statements and initialization code are generated. The target code is then merged with the common code elements to display the complete target code that meets the project's environment requirements.
[0259] Step 711: Detect and correct abnormalities in the target code.
[0260] In this step, the system checks the generated target code for anomalies, including syntax issues, data type mismatch issues, and code quality issues. If an anomaly is detected, an anomaly message and correction code are generated and displayed to the user. The user can then make corrections manually or automatically apply the correction code.
[0261] Step 712: Display the adapted code in the target programming language to the user.
[0262] In this step, the target code is syntactically converted according to the mapping rules between different programming languages to obtain the conversion code of the target programming language. Based on the programming rules of the target programming language, the code structure such as the code format and dependency declaration of the conversion code is adjusted, and the adaptation code adapted to the target programming language is generated and displayed to the user. For example, when converting Python code to C++ code, the grammatical structure of Python can be converted to an equivalent implementation of C++, and the programming standards and best practices of the C++ language are followed.
[0263] For example, if the user adds annotation content to the target code, the user behavior model can be updated to add user behavior characteristics of preference annotations. The mapping relationship between sentiment tendency categories and code styles can also be updated to improve the annotation richness of the code style corresponding to the sentiment tendency category that matches the input information, thereby more accurately meeting the user's personalized needs in subsequent code generation.
[0264] Figure 8 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present application. The electronic device may be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a personal digital assistant, a server, a smart home appliance, a car computer, etc. Figure 8 At the hardware level, the electronic device includes a processor 801, an internal bus 802, a network interface 803, a memory 804, and a non-volatile memory 805. Of course, it may also include hardware required for other services. The processor 801 reads the corresponding computer program from the non-volatile memory 805 into the memory 804 and then runs it, forming a code generation device at the logical level. Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0265] Figure 9 This is a block diagram of a code generation device according to an exemplary embodiment of the present application. Figure 9 The device may include: a data processing module 901, a sentiment analysis module 902 and a code generation module 903, wherein:
[0266] The data processing module 901 is used to obtain user input information and extract semantic features of the input information;
[0267] The sentiment analysis module 902 is configured to determine a sentiment tendency category matching the input information based on the semantic features;
[0268] The code generation module 903 is used to generate target code based on the semantic features and the emotional tendency category, and the coding style of the target code matches the emotional tendency category.
[0269] In one example, the input information includes current input information and historical input information; the data processing module 901, when used to extract semantic features of the input information, includes: performing feature extraction on the current input information to generate current semantic features corresponding to the current input information; obtaining historical semantic features corresponding to the historical input information; using the historical input information as context information of the current input information, superimposing the historical semantic features with the current semantic features to obtain the semantic features.
[0270] In one example, the sentiment analysis module 902, when used to determine the sentiment tendency category that matches the input information based on the semantic features, includes: determining the target code style of the target code based on the sentiment tendency category; generating initial code based on the semantic features; adjusting the code expression form of the initial code based on the target code style to generate the target code, and the code expression form includes at least one of the following: comments, logical structure.
[0271] In one example, the sentiment analysis module 902, when used to determine the sentiment tendency category matching the input information based on the semantic features, includes: assigning a probability value to each sentiment tendency category based on the semantic features; the sentiment analysis module 902, when used to determine the target code style of the target code based on the sentiment tendency category, includes: determining the sentiment probability values of different code styles based on the probability values of each sentiment tendency category; determining the user's preference probability values for the different code styles based on the user's historical behavior data; fusing the sentiment probability value and the preference probability value according to preset weights to generate final probability values of the different code styles; and using the code style with the highest final probability value as the target code style.
[0272] In one example, the sentiment analysis module 902, when used to determine the target code style of the target code according to the sentiment tendency category, includes: when the sentiment tendency category is a negative sentiment tendency, using the negative code style as the target code style; when the sentiment tendency category is a neutral sentiment tendency, using the neutral code style as the target code style; when the sentiment tendency category is a positive sentiment tendency, using the positive code style as the target code style; wherein the information richness of the auxiliary information of the negative code style, the neutral code style and the positive code style decreases.
[0273] In one example, the sentiment analysis module 902, when used to generate the initial code based on the semantic feature, includes: obtaining a set of candidate codes that match the semantic feature in a preset code library; the sentiment analysis module 902, when used to adjust the code expression of the initial code according to the target code style, includes: obtaining candidate codes that match the target code style in the candidate code set; and adjusting the candidate codes according to the target code style.
[0274] In one example, the sentiment analysis module 902, after being used to generate the initial code based on the semantic features, further includes: performing grammar detection on the initial code using grammar tree detection technology; in response to detecting that there are grammatical problems in the initial code, correcting the initial code to obtain a corrected initial code.
[0275] In one example, the code generation module 903 is also used to perform syntax conversion on the target code according to the mapping rules between different programming languages to obtain the conversion code of the target programming language; based on the programming rules of the target programming language, adjust the code structure of the conversion code to generate adaptation code, and the code structure includes at least one of the following: code format, dependency declaration.
[0276] In one example, the code generation module 903 is also used to perform anomaly detection on the target code to obtain a detection result; if the detection result is an abnormal result, generate abnormal prompt information and correction code for the target code; based on the correction instruction input by the user, use the correction code or the code carried in the correction instruction to correct the target code.
[0277] In one example, the code generation module 903, when used to perform anomaly detection on the target code, includes at least one of the following: detecting syntax problems of the target code; detecting data type matching problems of the target code; detecting code quality problems of the target code based on a preset rule engine; running the target code in a sandbox environment to detect abnormal running data of the target code; and performing unit testing on the target code.
[0278] In one example, the code generation module 903 is also used to generate common code elements based on the operating environment data of the target code and / or the code calling requirements of the input information in the operating environment, where the operating environment data includes at least one of the following: project environment resources, historical code data; and merge the common code elements with the target code.
[0279] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0280] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. Those of ordinary skill in the art can understand and implement it without paying any creative work.
[0281] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory including instructions. The instructions can be executed by a processor of a code generation device to implement any of the methods described in the above embodiments.
[0282] The non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc., and this application does not limit this.
[0283] In an exemplary embodiment, a computer program product including a computer program / instruction is further provided. The computer program / instruction can be executed by a processor of a code generation device to implement any of the methods described in the above embodiments.
[0284] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0285] Other embodiments of the present invention will readily occur to those skilled in the art after consideration of the specification and practice of the invention claimed herein. The present application is not limited to the precise construction described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present application is limited solely by the appended claims.
[0286] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A code generation method, characterized in that: The method comprises: Obtaining user input information and extracting semantic features of the input information; Determining, based on the semantic features, a sentiment tendency category that matches the input information; Based on the semantic features and the emotional tendency category, a target code is generated, and the coding style of the target code matches the emotional tendency category.
2. The method according to claim 1, characterized in that The input information includes current input information and historical input information; and extracting semantic features of the input information includes: Performing feature extraction on the current input information to generate current semantic features corresponding to the current input information; Obtaining historical semantic features corresponding to the historical input information; The historical input information is used as context information of the current input information, and the historical semantic features are superimposed on the current semantic features to obtain the semantic features.
3. The method according to claim 1, characterized in that The generating of the target code based on the semantic features and the emotional tendency category includes: Determining a target code style of the target code according to the emotional tendency category; generating an initial code based on the semantic features; According to the target code style, the code expression form of the initial code is adjusted to generate the target code, where the code expression form includes at least one of the following: comments and logical structure.
4. The method according to claim 3, characterized in that Determining the emotional tendency category matching the input information based on the semantic features includes: assigning probability values to each sentiment tendency category according to the semantic features; Determining the target code style of the target code according to the emotional tendency category includes: Determining the emotion probability values of different coding styles based on the probability values of the various emotion tendency categories; Determining, based on the user's historical behavior data, probability values of the user's preferences for the different coding styles; The emotion probability value and the preference probability value are combined according to a preset weight to generate a final probability value of the different coding styles; The coding style with the highest final probability value is used as the target coding style.
5. The method according to claim 3, characterized in that Determining the target code style of the target code according to the emotional tendency category includes: In the case where the sentiment tendency category is a negative sentiment tendency, the negative coding style is used as the target coding style; In the case where the emotional tendency category is neutral emotional tendency, the neutral coding style is used as the target coding style; In the case where the sentiment tendency category is positive sentiment tendency, the positive coding style is used as the target coding style; The information richness of the auxiliary information of the negative code style, the neutral code style and the positive code style decreases.
6. The method according to claim 3, characterized in that The generating of the initial code based on the semantic features includes: In a preset code library, a candidate code set matching the semantic feature is obtained; The adjusting the code representation of the initial code according to the target code style includes: Obtaining candidate codes that match the target code style from the candidate code set; The candidate code is adjusted according to the target code style.
7. The method according to claim 3, characterized in that After generating the initial code based on the semantic features, the method further includes: Performing syntax detection on the initial code using syntax tree detection technology; In response to detecting that the initial code has a grammatical problem, the initial code is corrected to obtain a corrected initial code.
8. The method according to claim 1, characterized in that The method further comprises: Performing syntax conversion on the target code according to mapping rules between different programming languages to obtain conversion code in the target programming language; Based on the programming rules of the target programming language, the code structure of the conversion code is adjusted to generate an adapted code, wherein the code structure includes at least one of the following: a code format and a dependency declaration.
9. The method according to claim 1, characterized in that The method further comprises: Performing anomaly detection on the target code to obtain a detection result; If the detection result is an abnormal result, generating abnormal prompt information and correction code for the target code; Based on the correction instruction input by the user, the target code is corrected using the correction code or the code carried in the correction instruction.
10. The method according to claim 9, characterized in that The performing abnormality detection on the target code includes at least one of the following: Detecting syntax problems in the target code; Detecting data type matching issues of the target code; Detecting code quality issues of the target code based on a preset rule engine; Running the target code in a sandbox environment to detect abnormal running data of the target code; Unit testing is performed on the target code.
11. The method according to claim 1, wherein The method further comprises: Generate a common code element according to the target code's runtime environment data and / or the code call requirements of the input information in the runtime environment, wherein the runtime environment data includes at least one of the following: project environment resources and historical code data; The universal code element and the target code are merged.
12. A code generating device, characterized in that: The device comprises: A data processing module is used to obtain user input information and extract semantic features of the input information; A sentiment analysis module, configured to determine a sentiment tendency category matching the input information based on the semantic features; The code generation module is used to generate target code based on the semantic features and the emotional tendency category, and the code style of the target code matches the emotional tendency category.
13. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 11 by running the executable instructions.
14. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
15. A computer program product having a computer program / instructions stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.