Method, device and equipment for generating auxiliary programming information
By performing reinforcement learning in the simulation system and adding image vector coding, the problem that it is difficult for general large language models to extract graphical information of programming problems is solved, and the accuracy of auxiliary programming information is improved.
Patent Information
- Application Number
- CN202411913822.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, it is difficult to accurately extract graphical information in programming questions, resulting in deviations in understanding programming questions and affecting the accuracy of programming auxiliary information.
By using pre-built auxiliary programming resources in the simulation system for reinforcement learning, an intermediate state sequence is obtained, and image vector encoding is added to the intermediate state sequence to form a target state sequence. When an end flag is generated in the target state sequence, it is inputted in combination with the preset prompt text to a large language model for analysis to generate auxiliary programming information.
By adding picture vector encoding, the visual representation of intermediate state sequences is enriched, which improves the accuracy of the understanding of programming problems by large language models, thereby improving the accuracy of auxiliary programming information.
Smart Images

Figure CN119987745A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device and equipment for generating auxiliary programming information. Background Art
[0002] With the rapid development of information technology, programming is no longer an exclusive skill for computer professionals. Artificial intelligence technology is gradually playing an important role in programming tutoring for teenagers. By collecting and analyzing a large amount of programming learning data, using machine learning and natural language technology, personalized learning suggestions and error diagnosis can be achieved. For example, the intelligent tutoring system can predict the difficulties that teenagers may encounter based on their learning history and provide solutions in advance.
[0003] In related technologies, programming questions and student answers can be combined as multiple rounds of dialogue and input into a general large language model. The ability of the general large language model can be used to assist students in programming, which can flexibly solve some programming problems with open questions. However, there are a large number of graphical problems in programming questions in the youth field. These graphical information usually contains complex visual logic. For the general large language model, it is difficult to accurately extract the graphical information, which leads to deviations in the understanding of programming questions and affects the accuracy of programming auxiliary information. Summary of the invention
[0004] In view of this, the present application provides a method and device for generating auxiliary programming information, the main purpose of which is to solve the problem in the prior art that it is difficult for general large language models to accurately extract graphical information from programming questions, resulting in deviations in the understanding of programming questions and affecting the accuracy of programming auxiliary information.
[0005] According to a first aspect of the present application, a method for generating auxiliary programming information is provided, comprising:
[0006] Use pre-built auxiliary programming resources to perform reinforcement learning solution operations in the simulation system to obtain intermediate state sequences;
[0007] Adding a picture vector code to the intermediate state sequence to obtain a target state sequence, wherein the picture vector code is obtained by processing the intermediate state sequence with at least one graphic conversion;
[0008] When an end flag is generated in the target state sequence, the target state sequence with picture vector encoding is combined with a preset prompt text and input into a large language model, so that the target state sequence is analyzed by the large language model to obtain auxiliary programming information with guiding function.
[0009] Furthermore, before performing the solution operation of reinforcement learning in the simulation system using the pre-built auxiliary programming resources, the method further includes:
[0010] Obtaining a standard operation sequence for auxiliary programming, wherein the standard operation sequence is obtained by serializing a building block code block input for programming;
[0011] Acquire a simulation configuration file of a programming topic, wherein the simulation configuration file includes a structured configuration file of the programming topic converted by a standard simulation module to obtain a code suitable for execution in a simulation environment;
[0012] According to the standard operation sequence and the simulation configuration file, auxiliary programming resources are constructed.
[0013] Furthermore, the standard operation sequence for obtaining auxiliary programming includes:
[0014] A block of code that receives programming input;
[0015] Converting the building block code blocks into standard codes executable by a computer system through standard coding processing;
[0016] Converting the standard code into an executable operation sequence through serialization processing to obtain a standard operation sequence for auxiliary programming;
[0017] The step of obtaining a simulation configuration file for a programming topic includes:
[0018] Obtaining a structured configuration file of a programming topic, wherein the structured configuration file includes at least position information and attribute information of each role in the programming topic;
[0019] The structured configuration file is converted into a code suitable for execution in a simulation environment through standard simulation processing to obtain a simulation configuration file of the programming topic.
[0020] Furthermore, the use of pre-built auxiliary programming resources to perform reinforcement learning solution operations in the simulation system to obtain an intermediate state sequence includes:
[0021] Performing a reinforcement learning solution operation in the simulation system using pre-built auxiliary programming resources to update the iteration resource according to the similarity distance between the simulation operation sequence and the standard operation sequence in the simulation system environment;
[0022] An operation sequence for replacing the standard operation sequence to complete the task set by the simulation configuration file is determined according to the updated iteration resources to obtain an intermediate state sequence.
[0023] Furthermore, the step of adding a graphic vector code to the intermediate state sequence to obtain a target state sequence includes:
[0024] Converting the intermediate state sequence into a visualized process picture through simulation configuration graphics;
[0025] Processing the visualized process image through graphic coding to obtain a picture vector code;
[0026] The graphic vector code is added to the end of the intermediate state sequence to obtain a target state sequence with the graphic vector code.
[0027] Furthermore, after adding the picture vector encoding to the intermediate state sequence to obtain the target state sequence, the method further includes:
[0028] If it is detected during the reinforcement learning process that the state value of the target state sequence meets the set condition, an end flag is generated in the target sequence.
[0029] Furthermore, when the end flag is generated in the target state sequence, the target state sequence having the picture vector encoding is combined with a preset prompt text and input into a large language model, so that the target state sequence is analyzed by the large language model to obtain auxiliary programming information with guiding effect, the method further includes:
[0030] The instructive programming auxiliary information is converted into instructive voice information through a voice generation model to provide programming guidance based on the instructive voice information.
[0031] According to a second aspect of the present application, there is provided a device for generating auxiliary programming information, comprising:
[0032] A solution unit, used for performing a solution operation of reinforcement learning in a simulation system using pre-built auxiliary programming resources to obtain an intermediate state sequence;
[0033] An adding unit, configured to add a picture vector code to the intermediate state sequence to obtain a target state sequence, wherein the picture vector code is obtained by processing the intermediate state sequence with at least one graphic conversion;
[0034] A generating unit is used for combining the target state sequence having the picture vector encoding with the preset prompt text and inputting the combined result into the large language model when an end flag is generated in the target state sequence, so as to analyze the target state sequence through the large language model and obtain auxiliary programming information with guiding function.
[0035] Furthermore, the device also includes:
[0036] A first acquisition unit is used to acquire a standard operation sequence of auxiliary programming before performing a solution operation of reinforcement learning in a simulation system using the pre-built auxiliary programming resources, wherein the standard operation sequence is obtained by serializing a building block code block input for programming;
[0037] A second acquisition unit is used to acquire a simulation configuration file of a programming topic, wherein the simulation configuration file includes a structured configuration file of the programming topic converted by a standard simulation module to obtain a code suitable for execution in a simulation environment;
[0038] A construction unit is used to construct auxiliary programming resources according to the standard operation sequence and the simulation configuration file.
[0039] Furthermore, the first acquisition unit is specifically configured to:
[0040] A block of code that receives programming input;
[0041] Converting the building block code blocks into standard codes executable by a computer system through standard coding processing;
[0042] Converting the standard code into an executable operation sequence through serialization processing to obtain a standard operation sequence for auxiliary programming;
[0043] The second acquisition unit is specifically used to:
[0044] Obtaining a structured configuration file of a programming topic, wherein the structured configuration file includes at least position information and attribute information of each role in the programming topic;
[0045] The structured configuration file is converted into a code suitable for execution in a simulation environment through standard simulation processing to obtain a simulation configuration file of the programming topic.
[0046] Furthermore, the solving unit is specifically used for:
[0047] Performing a reinforcement learning solution operation in the simulation system using pre-built auxiliary programming resources to update the iteration resource according to the similarity distance between the simulation operation sequence and the standard operation sequence in the simulation system environment;
[0048] An operation sequence for replacing the standard operation sequence to complete the task set by the simulation configuration file is determined according to the updated iteration resources to obtain an intermediate state sequence.
[0049] Furthermore, the adding unit is further used for:
[0050] Converting the intermediate state sequence into a visualized process picture through simulation configuration graphics;
[0051] Processing the visualized process image through graphic coding to obtain a picture vector code;
[0052] The graphic vector code is added to the end of the intermediate state sequence to obtain a target state sequence with the graphic vector code.
[0053] Furthermore, the device also includes:
[0054] The detection unit is used to add the picture vector code to the intermediate state sequence, and after obtaining the target state sequence, if it is detected in the process of reinforcement learning that the state value of the target state sequence meets the set conditions, an end flag is generated in the target sequence.
[0055] Furthermore, the device also includes:
[0056] A conversion unit is used to combine the target state sequence with picture vector encoding with preset prompt text and input it into a large language model when an end flag is generated in the target state sequence, so as to analyze the target state sequence through the large language model to obtain auxiliary programming information with guiding effect, and then convert the auxiliary programming information with guiding effect into guiding voice information through a voice generation model, so as to provide programming guidance according to the guiding voice information.
[0057] According to a third aspect of the present application, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in the first aspect when executing the computer program.
[0058] According to a fourth aspect of the present application, a readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0059] By means of the above technical scheme, the present application provides a method, device and equipment for generating auxiliary programming information. Compared with the method of using a general large language model for auxiliary programming in the prior art, the present application uses pre-built auxiliary programming resources to perform reinforcement learning solution operations in a simulation system to obtain an intermediate state sequence; add picture vector coding to the intermediate state sequence to obtain a target state sequence, and the picture vector coding is obtained by at least one graphic conversion process of the intermediate state sequence; when an end flag is generated in the target state sequence, the target state sequence with the picture vector coding is combined with a preset prompt text and input into the large language model, so as to analyze the target state sequence through the large language model to obtain auxiliary programming information with guiding effect. The whole process uses the intermediate state sequence obtained by reinforcement learning as the basis, adds picture vector coding to the intermediate state sequence, so as to enrich the general visual representation of the intermediate state sequence, and can make full use of the visual information processing capability of the large language model in the process of generating auxiliary programming information by the large language model, so that the auxiliary programming information has more accurate graphical information guidance, avoids the deviation of the understanding of the programming topic by the large language model, and improves the accuracy of the auxiliary programming information.
[0060] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0062] Figure 1 It is a flow chart of a method for generating auxiliary programming information provided by an embodiment of the present application;
[0063] Figure 2 It is a flowchart of another method for generating auxiliary programming information provided by an embodiment of the present application;
[0064] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step 201;
[0065] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step 202;
[0066] Figure 5 yes Figure 1 A schematic flow chart of a specific implementation of step 101;
[0067] Figure 6 yes Figure 1 A schematic flow chart of a specific implementation of step 102;
[0068] Figure 7 It is a flowchart of a method for generating auxiliary programming information provided by an embodiment of the present application;
[0069] Figure 8 It is a structural schematic diagram of a device for generating auxiliary programming information provided by an embodiment of the present application;
[0070] Fig. 9 It is a schematic diagram of the device structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] Now the content of the present invention will be discussed with reference to several exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and thus implement the content of the present invention, rather than implying any limitation on the scope of the present invention.
[0072] As used herein, the term "including" and variations thereof are to be interpreted as open-ended terms meaning "including but not limited to." The term "based on" is to be interpreted as "based, at least in part, on." The terms "one embodiment" and "an embodiment" are to be interpreted as "at least one embodiment." The term "another embodiment" is to be interpreted as "at least one other embodiment."
[0073] Related technologies can be implemented in the process of programming tutoring using the following methods:
[0074] In the first method, no matter what type of programming question it is, and no matter what specific problem the student's answer presents, the system can return a unique fixed standard answer, and the student can understand and modify the programming question based on this standard answer. However, this process gives a relatively rigid answer to programming questions with more flexible and open answers, and cannot provide accurate guidance based on the specific types of mistakes made by students. In this way, students may get closer to a solution with a slight modification, which makes it difficult to help improve students' programming level.
[0075] In the second method, the system will divide the programming questions into fixed answer types and flexible answer types through a classification model or manual judgment. The fixed answer type programming question system directly provides students with fixed answers, while the flexible type question system will be connected to the manual system. For programming questions connected to the manual system, after the students answer, the teacher will provide auxiliary answers and guidance. This process is fully compatible with flexible and multi-solution question exercises, but the quantity and quality requirements of human resources are relatively high, and it is difficult to support the explosive growth stage of the corresponding business.
[0076] The third method is to combine programming questions and students' answers as multiple rounds of dialogue, input them into the general language model, and use the ability of the general language model to assist students in programming, which can flexibly solve some programming problems with open questions. However, there are a large number of graphical problems in programming questions in the youth field. These graphical information usually contains complex visual logic. For the general language model, it is difficult to accurately extract the graphical information, which leads to deviations in the understanding of programming questions and affects the accuracy of programming auxiliary information.
[0077] In order to solve this problem, this embodiment provides a method for generating auxiliary programming information, such as Figure 1 As shown, the method comprises the following steps:
[0078] 101. Use pre-built auxiliary programming resources to perform reinforcement learning solution operations in the simulation system to obtain the intermediate state sequence.
[0079] Among them, the auxiliary programming resources at least include the operation sequence input by the user and the standard simulation configuration file converted from the programming questions. For programming learners, the operation sequence is a series of actions input when operating the programming software or platform, and records a series of operation steps performed by the user in the programming learning scenario. For example, in the context of robot control programming, it may include a sequence of operations with a sequence of controlling the robot to move forward, turn, grab objects, etc.
[0080] It is understandable that when the written program does not achieve the expected results, these operation sequences can also serve as key clues for troubleshooting problems. For example, when making a simple animation program, the character is expected to move along a specific trajectory, but in reality there is a deviation. By querying the operation sequence of the historical input, it can be found that the number of moving steps is set incorrectly or the logic of the conditional judgment is wrong, so that the program can be modified in a more targeted manner to continuously improve programming capabilities.
[0081] In some application scenarios, programming questions are sometimes relatively abstract. By refining and standardizing the requirements of programming questions to simulate the real programming environment and task requirements, abstract questions can be converted into specific rules, parameters and conditions, so that programmers can clearly understand how to build programs to meet the requirements of programming questions.
[0082] The standard simulation configuration file for a specific programming problem can be obtained through the following steps: first, analyze the programming problem to obtain a detailed description of the programming problem; then, use code to reconstruct the objects involved in the programming problem based on the detailed description of the programming problem to obtain the data structure of the object; if the programming problem involves different types of objects, a data structure can be designed for each object to represent it; divide the logic modules according to the goals and rules of the programming problem; write the code according to the data structure and logic modules of the object; and extract key parameters from the code and programming problem requirements to construct a standard simulation configuration file for the programming problem.
[0083] It is understandable that in order to construct a platform that interacts with the virtual environment, so that the auxiliary programming resources can perform various actions in the system and receive the content containing state information and reward information fed back by the environment, it is necessary to build the environment of the simulation system according to the task requirements configured by the programming topic. This process requires clarifying various parameters of the environment, defining the action space and state space according to the task characteristics, and the action space includes a set of actions that the agent can execute. For example, in the programming problem of robot path planning, the action set may include basic actions such as moving forward, moving backward, and moving left. The state space covers various state information that the agent needs to pay attention to in the environment, such as the current direction and position of the robot. The agent can choose the appropriate action based on the perception of the state. Then set the reward mechanism to guide the agent to learn. It can be designed according to the task goal configured by the programming topic. Taking the robot reaching the specified target location as an example, if the robot successfully moves to the target location, it can be given a higher positive reward. If the robot encounters an obstacle, it will be given a negative reward to encourage the agent to learn the correct behavior strategy and move towards the direction of completing the task. Further, the reinforcement learning algorithm can be selected according to the complexity of the programming problem, the characteristics of the state space and the action space. After the agent is initialized, it will interact continuously in the simulation system environment. In each time step, the agent will observe the current state of the environment, and then select an action from the action space according to the current strategy and execute it in the simulation system. For example, the robot decides to move forward one step, and the traffic agent adjusts the duration of the traffic light at a certain intersection. The environment of the simulation system will update its own state according to the result of the action, and feedback the corresponding reward and new environment state information to the agent. For example, after the robot moves forward, if it does not encounter obstacles and is closer to the target, it will receive a positive reward, and the environment state will be updated to the new position of the robot and the surrounding environment. Based on the rewards and new environment state information obtained, the agent uses the reinforcement learning algorithm to update its own strategy, so that the agent can make better decisions in subsequent interactions.
[0084] In other words, the purpose of solving the problem in reinforcement learning is to find an operation sequence that is as close as possible to the user input and can complete the task configured by the programming question. From the beginning of the interaction between the agent and the simulation system, the environmental state changes generated after each interaction constitute part of the intermediate state sequence. Recording these intermediate state sequences can obtain the state information of the agent in each time step, for example, recording the position, orientation, and distance of the robot from the target at each time step. By analyzing these intermediate state sequences, it is helpful to understand the learning trajectory of the agent, the effectiveness of the strategy, and the problems that may be encountered in the learning process, so as to optimize the reinforcement learning algorithm, adjust the simulation system parameters, or improve the configuration of the programming question, so as to better achieve the solution goal. In the process of reinforcement learning, by constantly trying and optimizing the strategy, the agent can learn which action combinations in the action space are more conducive to achieving the task goals under different environmental conditions. Correspondingly, the operation sequence corresponding to the action combination will be closer and closer to the ideal operation sequence given by the user.
[0085] 102. Add picture vector encoding to the intermediate state sequence to obtain a target state sequence.
[0086] In this embodiment, the picture vector is encoded as an intermediate state sequence obtained by at least one graphic conversion process. Since the intermediate state sequence contains different state information presented by the system environment at each time step in the solution operation process of reinforcement learning, these state information are arranged in chronological order to form a complete intermediate state sequence, which can reflect the dynamic changes of the system state during the entire solution operation process. Specifically, the abstract intermediate state sequence presented in the form of data can be converted into a visual picture through the conversion function of the simulation configuration graphic, and then the visual picture can be converted into a picture vector encoding.
[0087] Here, the target state sequence can be a structured JSON file, for example: {"reward value sequence": ["*****",], "distance from the end point": ["*****,], "image representation vector sequence": ["*******]}. The image vector encoding will be spliced to the tail of the image representation vector sequence in the intermediate state sequence to obtain the target state sequence.
[0088] It is understandable that picture vector coding has important applications in many aspects. For example, when comparing the effects of different reinforcement learning strategies, the similarity of the picture vector coding converted from the corresponding intermediate state sequence can be compared to judge the difference in system state changes caused by different strategies under similar initial conditions, and then evaluate the pros and cons of the strategies. When clustering the reinforcement learning process, similar system states can be classified into one category based on picture vector coding, which helps to discover typical state patterns at different stages, assist in optimizing learning algorithms or improving task configurations, etc. In addition, when building an image-based intelligent monitoring system or an automated analysis system, these picture vector codings can be easily integrated with other algorithm modules to achieve more efficient and automated analysis and management of the reinforcement learning process.
[0089] In actual operation, the intermediate state sequence is usually stored in a certain data structure, such as a list, array, or a custom structure. When adding the image vector code to the intermediate state sequence, it is only necessary to add a field or element to the data structure corresponding to each intermediate state to store the code. By adding the image vector code to the intermediate state sequence, it is possible to start from both the original state data and the image feature representation. On the one hand, numerical analysis is performed based on the original state data, for example, calculating the trend of state variables over time, analyzing the impact of different actions on the state, etc.; on the other hand, the image vector code is used to mine information from the perspective of image features. For example, similar intermediate state images are grouped together based on vector codes through cluster analysis to discover typical state patterns or abnormal states in the reinforcement learning process, thereby gaining a deeper understanding of the effectiveness of the learning path and strategy of the intelligent agent.
[0090] 103. When an end flag is generated in the target state sequence, the target state sequence having the image vector encoding is combined with a preset prompt text and input into a large language model, so that the target state sequence is analyzed by the large language model to obtain auxiliary programming information with guiding function.
[0091] In this embodiment, the end flag is a specific identifier used to indicate that the entire reinforcement learning process has reached a certain preset end condition, where the end condition is usually set according to the task requirements configured by the programming topic. For example, it is set to the distance d from the end target, and the end mark in the target state sequence is d<=0. When d<=0 in the target state sequence, the reinforcement learning process is judged to be over. The end flag represents that the entire reinforcement learning task has entered the end stage in the current round or instance. If the end flag is generated in the target state sequence, the target state sequence with picture vector encoding can be extracted and connected with other modules to mine valuable information in the target state sequence and provide a reference for auxiliary programming.
[0092] It can be understood that in order to accurately prompt the large language model to pay attention to information focus, analysis ideas and expected results when analyzing the target state sequence, the preset prompt text serves as the analysis framework and direction provided by the large prediction model, which enables the large language model to more accurately extract useful information from the target state sequence and generate auxiliary programming information that meets the requirements.
[0093] Specifically, the target state sequence and the preset prompt text are usually spliced into an input content according to certain format requirements. The input content can be a simple text splicing form, for example, the preset prompt text is written in front, followed by the target state sequence data serialized in a specific format, or it can be more carefully structured according to the input interface specification required by the large language model, for example, the preset prompt text is used as a field, and the target state sequence is used as another field, together forming a JSON object of an input request and inputting it into the large language model. After the above splicing process, the large language model can rely on its powerful natural language processing ability and the knowledge understanding and reasoning ability after learning a large amount of text data to process the received input content containing the target state sequence and the preset prompt text, so as to comprehensively analyze the target state sequence and generate corresponding auxiliary programming information according to the requirements of the preset prompt text. This information can provide guiding information for programming questions, which can be guiding information on functional integrity, for example, whether the core required by the programming question is realized, whether the auxiliary functions are complete, and provide functional expansion ideas, and can also be guidance on code structure and logic, for example, whether the code logic is accurate and whether the code has clear module division.
[0094] For example, the input content formed by the concatenation of the target state sequence and the preset prompt text can be expressed as follows: Preset prompt text: You are a children's programming teacher. The student's answer to the current question is ****. The image information generated by the student's problem-solving process is ***** (the image vector encoding in the above text). The state of the intermediate process of solving the problem is: target state sequence. Please generate the corresponding guidance text ***** based on the target state sequence in solving the problem.
[0095] The method for generating auxiliary programming information provided by the embodiment of the present application is compared with the method of using a general large language model for auxiliary programming in the current prior art. The present application uses pre-built auxiliary programming resources to perform reinforcement learning solution operations in a simulation system to obtain an intermediate state sequence; adds picture vector coding to the intermediate state sequence to obtain a target state sequence, and the picture vector coding is obtained by at least one graphic conversion process of the intermediate state sequence; when an end flag is generated in the target state sequence, the target state sequence with the picture vector coding is combined with a preset prompt text and input into the large language model, so as to analyze the target state sequence through the large language model to obtain auxiliary programming information with guiding effect. The whole process uses the intermediate state sequence obtained by reinforcement learning as the basis, adds picture vector coding to the intermediate state sequence, so as to enrich the general visual representation of the intermediate state sequence, and can make full use of the visual information processing capability of the large language model in the process of generating auxiliary programming information by the large language model, so that the auxiliary programming information has more accurate graphical information guidance, avoids the deviation of the understanding of the programming topic by the large language model, and improves the accuracy of the auxiliary programming information.
[0096] In actual application scenarios, considering that programming problems often have different levels of complexity, auxiliary programming resources can be pre-built to clearly understand the goals, environment, and rules of programming tasks. Figure 2 As shown, before step 101, the method further includes the following steps:
[0097] 201. Obtain a standard operation sequence for auxiliary programming.
[0098] 202. Obtain a simulation configuration file for the programming question.
[0099] 203. Construct auxiliary programming resources according to the standard operation sequence and the simulation configuration file.
[0100] The standard operation sequence is obtained by processing the input building block code blocks through standard serialization. Figure 3 As shown, step 201 includes the following steps:
[0101] 301. Receive a building block code block for programming input.
[0102] 302. Convert the building block code block into a standard code executable by a computer system through standard coding processing.
[0103] 303. Convert the standard code into an executable operation sequence through serialization processing to obtain a standard operation sequence for auxiliary programming.
[0104] In a visual programming environment, programming tools can usually be implemented by dragging and splicing building block code blocks. These building block code blocks each represent different programming instructions or functions, such as the "if, then" building block code for control flow, the "move a few steps" building block code for character movement, etc. The building block code input for programming can be combined and arranged in a certain logical order to obtain an operation sequence to achieve the operation goal. In order to facilitate subsequent storage, analysis and reuse in different scenarios, these operation sequences need to be serialized to convert the operation sequence in the form of visual building block code splicing into a unified and recognizable data format, so that the serialized operation sequence can be easily read and parsed by the computer program.
[0105] The simulation configuration file includes a structured configuration file of a programming topic, which is converted into a code suitable for execution in a simulation environment through a standard simulation module. Figure 4 As shown, step 202 includes the following steps:
[0106] 401. Get a structured configuration file for programming questions.
[0107] 402. Convert the structured configuration file into a code suitable for execution in a simulation environment through standard simulation processing to obtain a simulation configuration file of the programming problem.
[0108] In this embodiment, the structured configuration file at least includes the location information and attribute information of each role in the programming problem. Here, the structured configuration file of the programming problem is usually set by the problem setter or programming teacher according to the requirements of the programming task, and usually covers many aspects of content, which is equivalent to the information structured matching of the programming problem, for example, the coordinates of each role in the figure, the basic attributes of each role and other problem information, specifically including the environment related information of the programming task, the target description of the programming task and the input and output requirements. Through the processing of the standard simulation module, the structured configuration file can be converted into a code that can be directly executed in the simulation environment. Here, the standard simulation module will use the corresponding programming library or framework to create the corresponding virtual environment according to the environment parameters described in the configuration information. The simulation configuration file finally generated is the overall code file or code module set that integrates the environment construction code, the rule implementation code and the input and output processing code. Through the code module set, the programming problem can be accurately simulated and run in the simulation environment, which is convenient for learners to try to solve the programming problem.
[0109] Correspondingly, such as Figure 5 As shown, step 101 includes the following steps:
[0110] 501. Use pre-built auxiliary programming resources to perform reinforcement learning solution operations in the simulation system to update iterative resources according to the similarity distance between the simulation operation sequence and the standard operation sequence in the simulation system environment.
[0111] 502. Determine, according to the updated iteration resources, an operation sequence for replacing the standard operation sequence to complete the task set in the simulation configuration file, and obtain an intermediate state sequence.
[0112] It is understandable that in the initial stage of reinforcement learning, the initial strategy of the agent is relatively random. Then, according to the key parameters of the reinforcement learning algorithm, such as learning rate, discount factor and exploration rate, preparations are made for the subsequent iterative learning process. Then, in the process of solving the reinforcement learning problem, the agent continuously interacts with the environment of the simulation system, selects and executes corresponding actions according to the current state of the environment and its own strategy. As time goes by, the actions executed in sequence constitute a simulated operation sequence, which is used to reflect the behavior trajectory of the agent when trying to complete the task.
[0113] Furthermore, in order to measure the degree of difference between the simulated operation sequence and the standard operation sequence, the method based on the vector space model can be used to convert the operation sequence into a vector representation, and then the similarity distance calculated by the distance metric between the vectors is used to represent the similarity. The smaller the similarity distance, the more similar the two are. Then, according to the calculated similarity distance, corresponding feedback is provided to the agent to prompt it to adjust its strategy. For example, if the similarity distance is large, it means that the current agent's operation sequence is significantly different from the standard operation sequence, which means that the agent's strategy is not reasonable enough. At this time, the simulation system will give the agent a relatively low reward or negative reward to encourage it to explore other possible action choices, so that the subsequent simulated operation sequence is closer to the standard operation sequence. On the contrary, if the similarity distance is small, it means that the agent's behavior is better and closer to the ideal operation mode. A higher positive reward is given to strengthen the current strategy and increase the probability of continuing to choose similar actions in the future.
[0114] In this embodiment, through the continuous attempts and learning of the agent in the simulation system, a relatively stable operation sequence that can effectively complete the task set by the simulation configuration file is gradually determined. This operation sequence is determined by adjusting the action selection according to the environmental feedback in multiple iterations. In the process of the agent performing the task according to the determined operation sequence, the state of the simulation system at each time step is recorded, and these state sets arranged in chronological order constitute the intermediate state sequence.
[0115] In actual application scenarios, considering that the state of the model loading process in different game scenarios will change, specifically, Figure 6As shown, step 102 includes the following steps:
[0116] 601. Convert the intermediate state sequence into a visualized process picture through simulation configuration graphics.
[0117] 602. Process the visualized process image through graphic coding to obtain image vector coding.
[0118] 603. Add the graphic vector code to the end of the intermediate state sequence to obtain a target state sequence with the graphic vector code.
[0119] Specifically, in the process of converting the intermediate state sequence into a visualized process picture through simulation configuration graphics, it can be implemented according to the pre-set simulation configuration rules and graphics drawing logic, parsing the data in the intermediate state sequence, and combining these elements on the canvas in appropriate proportions, color matching and layout methods, and finally generating a visualized picture that can show the system status at that moment.
[0120] Specifically, in the process of obtaining the image vector code by processing the visualized process image through graphic coding, it can be realized based on the convolutional neural network. First, the image input value is pre-trained in the convolutional neural network. The network has been pre-trained with a large amount of image data and has learned different feature representations of the image. For example, it can recognize the underlying features such as edges, textures, shapes, etc. in the image, and gradually abstract higher-level semantic features. After the visualized image has undergone multi-layer convolution, pooling and other operations of the network, a fixed-length vector code is finally output. This vector code represents the feature representation of the visualized image and condenses the key information in the image, so that the computer can perform subsequent comparison, classification, analysis and other operations based on these vector codes.
[0121] In actual application scenarios, whether the reinforcement learning process has successfully completed the task or reached a critical stage at which the current learning round can be ended can be considered from multiple aspects, such as task completion indicators, performance optimization thresholds, resource constraints or time constraints, etc. Further, before step 102, the method further includes the following steps:
[0122] If it is detected during the reinforcement learning process that the state value of the target state sequence meets the set condition, an end flag is generated in the target sequence.
[0123] It is understandable that at each time step or key stage of reinforcement learning, the system will detect the state value in the target state sequence to determine whether it meets the pre-set conditions. This usually requires embedding the corresponding judgment logic in the code, for example, using conditional judgment statements to check whether the values of each related state variable reach the corresponding threshold or conform to a specific logical relationship. Further, the end flag is generated based on whether the state value of the target state sequence meets the set conditions. As an effective task end judgment mechanism, the end flag plays an indispensable role in the operation, evaluation and continuous optimization of the entire reinforcement learning system.
[0124] In actual application scenarios, in order to improve the flexibility of the auxiliary programming information in providing guidance, further, after step 103, the method further includes the following steps:
[0125] The instructive programming auxiliary information is converted into instructive voice information through a voice generation model to provide programming guidance based on the instructive voice information.
[0126] In this embodiment, the auxiliary programming information is usually extracted through a multi-faceted analysis of programming problems, an evaluation of programming results, or a summary based on past programming experience, with the purpose of helping programmers better understand programming problems, improve codes, and improve programming skills. The speech generation model here is an artificial intelligence model built based on technologies such as deep learning. By learning a large amount of text and speech corresponding data, it can provide the ability to convert text information into natural and fluent speech. Specifically, the programming auxiliary information is input into the speech generation model, and the speech generation model performs text preprocessing on the programming auxiliary information, such as word segmentation, part-of-speech tagging, etc., in order to better understand the grammatical structure and semantic information of the text in the programming auxiliary information. Then, based on the speech synthesis rules and parameters learned by the speech generation model, the corresponding pronunciation mode, pitch, length and other acoustic features are determined for each word, and these features are combined through a synthesis algorithm to generate a continuous speech waveform signal, which is finally output as a playable guidance voice information.
[0127] It should be noted that different speech generation models may differ in specific conversion details and effects, but the overall goal is to convert the input programming auxiliary information into speech form as accurately and naturally as possible.
[0128] Specifically in the actual application scenario, the generation process of auxiliary programming information is as follows Figure 7As shown, on the one hand, a user input building block code block m1 is received, and the building block code block m1 can be converted into a standard code o1 executable by the computer system through the standardization module, and then the standard code o1 is converted into an executable operation sequence c1 through the serialization module; on the other hand, a configuration file k1 of a programming question is received, and the configuration file k1 is converted into a standard simulation configuration file f1 acceptable to the simulation environment through the simulation module; then, the computer simulation system receives the operation sequence c1 and the standard simulation configuration file f1, and combines the operation sequence c1 with the standard simulation configuration file f1 to perform a reinforcement learning solution operation in the computer simulation system. The solution purpose is to find a target operation sequence as close to the operation sequence c1 as possible and capable of completing the task configured by the quasi-simulation configuration file f1, and record the reinforcement learning. The intermediate state sequence s1 obtained by solving the learning process is further converted into a visualized process image p1 through a simulation configuration graphic conversion module, and then the process image p1 is processed into an image vector code e1 through a graphic encoding module, and the image vector code e1 is merged into the intermediate state sequence s1 to obtain a target state sequence s2, and the target state sequence s2 is analyzed and judged to determine whether the reinforcement learning solution process is completed. If not, the above-mentioned computer simulation system is repeated to perform the reinforcement learning solution process. If so, the marked state sequence s2 is combined with a preset prompt text as the input of the large language model, and the intermediate state is comprehensively analyzed through the large language model to obtain a guidance text, which is further converted into a guidance speech through a text-to-speech model.
[0129] Further, as Figure 1-Figure 6 The specific implementation of the method, the embodiment of the present application provides a device for generating auxiliary programming information, such as Figure 8 As shown, the device includes: a solving unit 71, an adding unit 72, and a generating unit 73.
[0130] A solving unit 71, used to perform a solving operation of reinforcement learning in a simulation system using pre-built auxiliary programming resources to obtain an intermediate state sequence;
[0131] An adding unit 72 is used to add a picture vector code to the intermediate state sequence to obtain a target state sequence, wherein the picture vector code is obtained by processing the intermediate state sequence through at least one graphic conversion;
[0132] The generating unit 73 is used for combining the target state sequence having the picture vector encoding with the preset prompt text and inputting them into the large language model when an end flag is generated in the target state sequence, so as to analyze the target state sequence through the large language model and obtain auxiliary programming information with guiding function.
[0133] The device for generating auxiliary programming information provided by the embodiment of the present invention is compared with the method of using a general large language model for auxiliary programming in the current prior art. The present application uses pre-built auxiliary programming resources to perform reinforcement learning solution operations in a simulation system to obtain an intermediate state sequence; adds picture vector coding to the intermediate state sequence to obtain a target state sequence, and the picture vector coding is obtained by processing the intermediate state sequence through at least one graphic conversion; when an end flag is generated in the target state sequence, the target state sequence with the picture vector coding is combined with a preset prompt text and input into the large language model, so as to analyze the target state sequence through the large language model to obtain auxiliary programming information with guiding effect. The whole process uses the intermediate state sequence obtained by reinforcement learning as the basis, adds picture vector coding to the intermediate state sequence, so as to enrich the general visual representation of the intermediate state sequence, and can make full use of the visual information processing capability of the large language model in the process of generating auxiliary programming information by the large language model, so that the auxiliary programming information has more accurate graphical information guidance, avoids the deviation of the understanding of the programming topic by the large language model, and improves the accuracy of the auxiliary programming information.
[0134] In an actual application scenario, the device further includes:
[0135] A first acquisition unit is used to acquire a standard operation sequence of auxiliary programming before performing a solution operation of reinforcement learning in a simulation system using the pre-built auxiliary programming resources, wherein the standard operation sequence is obtained by serializing a building block code block input for programming;
[0136] A second acquisition unit is used to acquire a simulation configuration file of a programming topic, wherein the simulation configuration file includes a structured configuration file of the programming topic converted by a standard simulation module to obtain a code suitable for execution in a simulation environment;
[0137] A construction unit is used to construct auxiliary programming resources according to the standard operation sequence and the simulation configuration file.
[0138] In an actual application scenario, the first acquisition unit is specifically used to:
[0139] A block of code that receives programming input;
[0140] Converting the building block code blocks into standard codes executable by a computer system through standard coding processing;
[0141] Converting the standard code into an executable operation sequence through serialization processing to obtain a standard operation sequence for auxiliary programming;
[0142] The second acquisition unit is specifically used to:
[0143] Obtaining a structured configuration file of a programming topic, wherein the structured configuration file includes at least position information and attribute information of each role in the programming topic;
[0144] The structured configuration file is converted into a code suitable for execution in a simulation environment through standard simulation processing to obtain a simulation configuration file of the programming topic.
[0145] In actual application scenarios, the solving unit is specifically used for:
[0146] Performing a reinforcement learning solution operation in the simulation system using pre-built auxiliary programming resources to update the iteration resource according to the similarity distance between the simulation operation sequence and the standard operation sequence in the simulation system environment;
[0147] An operation sequence for replacing the standard operation sequence to complete the task set by the simulation configuration file is determined according to the updated iteration resources to obtain an intermediate state sequence.
[0148] In actual application scenarios, the adding unit is further used for:
[0149] Converting the intermediate state sequence into a visualized process picture through simulation configuration graphics;
[0150] Processing the visualized process image through graphic coding to obtain a picture vector code;
[0151] The graphic vector code is added to the end of the intermediate state sequence to obtain a target state sequence with the graphic vector code.
[0152] In an actual application scenario, the device further includes:
[0153] The detection unit is used to add the picture vector code to the intermediate state sequence, and after obtaining the target state sequence, if it is detected in the process of reinforcement learning that the state value of the target state sequence meets the set conditions, an end flag is generated in the target sequence.
[0154] In an actual application scenario, the device further includes:
[0155] A conversion unit is used to combine the target state sequence with picture vector encoding with preset prompt text and input it into a large language model when an end flag is generated in the target state sequence, so as to analyze the target state sequence through the large language model to obtain auxiliary programming information with guiding effect, and then convert the auxiliary programming information with guiding effect into guiding voice information through a voice generation model, so as to provide programming guidance according to the guiding voice information.
[0156] It should be noted that for other corresponding descriptions of the functional units involved in the device for generating auxiliary programming information provided in this embodiment, reference can be made to Figure 1-Figure 6 The corresponding description in will not be repeated here.
[0157] Based on the above Figure 1-Figure 6 The method shown in the embodiment of the present application accordingly provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned Figure 1-Figure 6 The method for generating auxiliary programming information is shown.
[0158] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation scenario of the present application.
[0159] Based on the above Figure 1-Figure 6 The method shown, and Figure 8 In order to achieve the above-mentioned purpose, the embodiment of the present application also provides a physical device for assisting the generation of programming information, which can be a computer, a smart phone, a tablet computer, a smart watch, a server, or a network device, etc. The physical device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figure 1-Figure 6 The method for generating auxiliary programming information is shown.
[0160] Optionally, the physical device may also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0161] In an exemplary embodiment, see Fig. 9 The physical device includes a communication bus, a processor, a memory and a communication interface, and may also include an input / output interface and a display device, wherein each functional unit can communicate with each other through the bus. The memory stores a computer program, and the processor is used to execute the program stored in the memory and execute the method for generating auxiliary programming information in the above embodiment.
[0162] Those skilled in the art will appreciate that the physical device structure for assisting in the generation of programming information provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different arrangements of components.
[0163] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device for generating and processing the auxiliary programming information, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to realize the communication between the components inside the storage medium, and the communication with other hardware and software in the information processing physical device.
[0164] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform, or by hardware. By applying the technical solution of the present application, compared with the current existing methods, the present application uses the intermediate state sequence obtained by reinforcement learning as the basis, and adds picture vector encoding to the intermediate state sequence to enrich the general visual representation of the intermediate state sequence. In the process of generating auxiliary programming information by the large language model, the visual information processing capability of the large language model can be fully utilized, so that the auxiliary programming information has more accurate graphical information guidance, avoiding deviations in the understanding of programming questions by the large language model, and improving the accuracy of the auxiliary programming information.
[0165] Those skilled in the art will appreciate that the accompanying drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily necessary for implementing the present application. Those skilled in the art will appreciate that the modules in the devices in the implementation scenario can be distributed in the devices of the implementation scenario according to the description of the implementation scenario, or can be changed accordingly and located in one or more devices different from the present implementation scenario. The modules of the above-mentioned implementation scenario can be combined into one module, or can be further split into multiple submodules.
[0166] The above serial numbers of this application are only for description and do not represent the advantages and disadvantages of the implementation scenarios. The above disclosure is only a few specific implementation scenarios of this application, but this application is not limited to them, and any changes that can be thought of by technicians in this field should fall within the scope of protection of this application.
Claims
1. A method for generating auxiliary programming information, characterized in that: include: Use pre-built auxiliary programming resources to perform reinforcement learning solution operations in the simulation system to obtain intermediate state sequences; Adding a picture vector code to the intermediate state sequence to obtain a target state sequence, wherein the picture vector code is obtained by processing the intermediate state sequence through at least one graphic conversion; When an end flag is generated in the target state sequence, the target state sequence with picture vector encoding is combined with a preset prompt text and input into a large language model, so that the target state sequence is analyzed by the large language model to obtain auxiliary programming information with guiding function.
2. The method according to claim 1, characterized in that Before performing the solution operation of reinforcement learning in the simulation system using the pre-built auxiliary programming resources, the method further includes: Obtaining a standard operation sequence for auxiliary programming, wherein the standard operation sequence is obtained by subjecting a programming input building block code block to standard serialization processing; Acquire a simulation configuration file of a programming topic, wherein the simulation configuration file includes a structured configuration file of the programming topic converted by a standard simulation module to obtain a code suitable for execution in a simulation environment; According to the standard operation sequence and the simulation configuration file, auxiliary programming resources are constructed.
3. The method according to claim 2, characterized in that The standard operation sequence for obtaining auxiliary programming includes: A block of code that receives programming input; Converting the building block code blocks into standard codes executable by a computer system through standard coding processing; Converting the standard code into an executable operation sequence through serialization processing to obtain a standard operation sequence for auxiliary programming; The step of obtaining a simulation configuration file for a programming topic includes: Obtaining a structured configuration file of a programming topic, wherein the structured configuration file includes at least position information and attribute information of each role in the programming topic; The structured configuration file is converted into a code suitable for execution in a simulation environment through standard simulation processing to obtain a simulation configuration file of the programming topic.
4. The method according to claim 2, characterized in that: The use of pre-built auxiliary programming resources to perform reinforcement learning solution operations in the simulation system to obtain an intermediate state sequence includes: Performing a reinforcement learning solution operation in the simulation system using pre-built auxiliary programming resources to update the iteration resource according to the similarity distance between the simulation operation sequence and the standard operation sequence in the simulation system environment; An operation sequence for replacing the standard operation sequence to complete the task set by the simulation configuration file is determined according to the updated iteration resources to obtain an intermediate state sequence.
5. The method according to claim 1, characterized in that Adding the graphic vector code to the intermediate state sequence to obtain the target state sequence includes: Converting the intermediate state sequence into a visualized process picture through simulation configuration graphics; Processing the visualized process image through graphic coding to obtain a picture vector code; The graphic vector code is added to the end of the intermediate state sequence to obtain a target state sequence with the graphic vector code.
6. The method according to claim 1, characterized in that After adding the picture vector code to the intermediate state sequence to obtain the target state sequence, the method further includes: If it is detected during the reinforcement learning process that the state value of the target state sequence meets the set condition, an end flag is generated in the target sequence.
7. The method according to any one of claims 1 to 6, characterized in that When the end flag is generated in the target state sequence, the target state sequence having the picture vector encoding is combined with the preset prompt text and input into the large language model, so as to analyze the target state sequence through the large language model to obtain the auxiliary programming information with guiding effect, the method further includes: The instructive programming auxiliary information is converted into instructive voice information through a voice generation model to provide programming guidance based on the instructive voice information.
8. A device for generating auxiliary programming information, characterized in that: include: A solution unit, used for performing a solution operation of reinforcement learning in a simulation system using pre-built auxiliary programming resources to obtain an intermediate state sequence; An adding unit, configured to add a picture vector code to the intermediate state sequence to obtain a target state sequence, wherein the picture vector code is obtained by processing the intermediate state sequence with at least one graphic conversion; A generating unit is used for combining the target state sequence having the picture vector encoding with the preset prompt text and inputting the combined result into the large language model when an end flag is generated in the target state sequence, so as to analyze the target state sequence through the large language model and obtain auxiliary programming information with guiding function.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for generating auxiliary programming information according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating auxiliary programming information according to any one of claims 1 to 7 are implemented.