Intelligent teaching method and system, electronic equipment, storage medium and product

By determining the user's fingertip position and algorithm type through an intelligent teaching system, and allowing the user to set logic block parameters independently, the disconnect between theoretical and practical learning in computer algorithm courses has been resolved, thereby improving teaching efficiency and quality.

CN122072934APending Publication Date: 2026-05-22CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD
Filing Date
2024-11-21
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

In traditional teaching methods, there is a significant disconnect between the theoretical and practical learning of computer algorithms, making it difficult to effectively integrate them and thus limiting teaching efficiency and quality.

Method used

The intelligent teaching system uses a binocular camera to acquire images of the user's hand, determine the fingertip position, identify the text region to be detected and determine the target algorithm type, extract relevant logic blocks from the logic block library, and allow the user to set parameters to achieve the integration of theory and practice.

Benefits of technology

This approach achieves a close integration of theoretical and practical learning, improves teaching efficiency and quality, and meets personalized learning needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122072934A_ABST
    Figure CN122072934A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent teaching method and system, electronic equipment, a computer storage medium and a computer program product, and the method is applied to the technical field of computers, and comprises the steps: obtaining a binocular image containing a hand of a user, and determining a fingertip position of the hand of the user based on the binocular image; obtaining a to-be-detected text area in a set range of the fingertip position, and determining a target algorithm type corresponding to the to-be-detected text area; all logic blocks related to the target algorithm type are extracted from a logic block library, and when parameter setting trigger operation for the target logic block in all the logic blocks is detected, the target logic block subjected to parameter setting is obtained; and fusing the target logic block after parameter setting and the logic block without parameter setting to obtain a fusion result, and displaying the fusion result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an intelligent teaching method, system, electronic device, computer storage medium, and computer program product. Background Technology

[0002] Currently, traditional teaching methods are teacher-centered, with students passively receiving education within a relatively uniform teaching pace and learning resources. However, students' intelligence, interests, and learning abilities vary, and this uniform traditional teaching approach cannot meet the needs of all students. The emergence of intelligent teaching methods can provide students with tailored learning resources and learning paths based on their individual needs, enabling students to learn more efficiently and autonomously.

[0003] In related technologies, for courses that combine theoretical teaching and practical learning, the current teaching method still adopts the approach of lecturing on theoretical logic in the classroom and conducting practical learning after class. This leads to a disconnect between teaching theory and practical learning, which to some extent creates a teaching bottleneck. Summary of the Invention

[0004] This application provides an intelligent teaching method, system, electronic device, computer storage medium, and computer program product.

[0005] The technical solution of this application is implemented as follows:

[0006] This application provides an intelligent teaching method applied to an intelligent teaching system, the method comprising:

[0007] Acquire a binocular image containing the user's hand, and determine the fingertip position of the user's hand based on the binocular image;

[0008] Obtain the text region to be detected within the set range of the fingertip position, and determine the target algorithm type corresponding to the text region to be detected;

[0009] Extract each logic block related to the target algorithm type from the logic block library, and when a parameter setting trigger operation is detected for the target logic block in each logic block, obtain the target logic block after parameter setting.

[0010] The target logic block with parameters set and the logic block without parameters set are fused together to obtain a fusion result, which is then displayed.

[0011] This application provides an intelligent teaching system, the system comprising:

[0012] An image perception module is used to acquire a binocular image containing the user's hand, and to determine the position of the fingertips of the user's hand based on the binocular image;

[0013] The text recognition module is used to acquire the text region to be detected within a set range of the fingertip position, and to determine the target algorithm type corresponding to the text region to be detected;

[0014] The logic construction module is used to extract various logic blocks related to the target algorithm type from the logic block library, and when a parameter setting trigger operation is detected for the target logic block in each logic block, the target logic block after parameter setting is obtained; the target logic block after parameter setting and the logic block without parameter setting are fused to obtain a fusion result.

[0015] The visualization module is used to display the fusion results.

[0016] This application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the intelligent teaching method provided by one or more of the foregoing technical solutions.

[0017] This application provides a computer storage medium storing a computer program; when the computer program is executed, it can implement the intelligent teaching method provided by one or more of the aforementioned technical solutions.

[0018] This application provides a computer program product, including a computer program that, when executed by a processor, implements the intelligent teaching method provided by one or more of the aforementioned technical solutions.

[0019] This application provides an intelligent teaching method, system, electronic device, computer storage medium, and computer program product. The method includes: acquiring a binocular image containing a user's hand; determining the fingertip position of the user's hand based on the binocular image; acquiring a text region to be detected within a set range of the fingertip position, and determining the target algorithm type corresponding to the text region to be detected; extracting various logic blocks related to the target algorithm type from a logic block library; when a parameter setting trigger operation is detected for the target logic block in each logic block, acquiring the target logic block after parameter setting; fusing the target logic block after parameter setting and the logic block without parameter setting to obtain a fusion result and displaying the fusion result.

[0020] As can be seen, the intelligent teaching method of this application first determines the position of the user's fingertip, then determines the target algorithm type corresponding to the text region to be detected pointed to by the fingertip, and then extracts each logic block related to the target algorithm type. The user sets the parameters of each logic block to realize the corresponding algorithm function. The whole process can be carried out while the user learns the theoretical logic. In this way, it can ensure that the user completes the practical learning of the relevant algorithm while mastering the theoretical logic. That is, it can realize the integration between teaching theory and practical learning, which is conducive to improving teaching efficiency and quality and ensuring teaching effect. Attached Figure Description

[0021] Figure 1 A flowchart illustrating an intelligent teaching method provided in this application embodiment;

[0022] Figure 2 This is a schematic diagram of the composition structure of the intelligent teaching system provided in the embodiments of this application;

[0023] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0024] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present application and are not intended to limit the present application. Furthermore, the embodiments provided below are some embodiments for implementing the present application, and not all embodiments for implementing the present application. Unless otherwise specified, the technical solutions described in the embodiments of the present application can be implemented in any combination.

[0025] It should be noted that, in the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or system that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or system. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the system; for example, a unit may be a portion of circuitry, a portion of a processor, a portion of a program or software, etc.) in the method or system that includes that element.

[0026] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, I and / or J can represent three cases: I alone, I and J simultaneously, and J alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of elements. For example, including at least one of I, J, and R can mean including any one or more elements selected from the set consisting of I, J, and R.

[0027] For example, the intelligent teaching method provided in this application includes a series of steps, but the intelligent teaching method provided in this application is not limited to the steps described. Similarly, the intelligent teaching system provided in this application includes a series of modules, but the intelligent teaching system provided in this application is not limited to the modules explicitly described, but may also include modules that need to be set up for obtaining relevant information or processing based on information.

[0028] Currently, intelligent teaching methods for computer algorithm courses are relatively scarce. Besides mastering theoretical logic, the most crucial aspect of computer algorithm courses is implementing, verifying, and improving algorithms—this is the core of the course and a vital way to enhance algorithmic skills. However, current teaching methods still rely on classroom lectures to impart theoretical logic, followed by practical learning through computer lab sessions. This creates a significant disconnect between the two types of instruction, hindering the integration of theoretical and practical learning and thus creating a teaching bottleneck.

[0029] To address the aforementioned technical problems, the following embodiments are proposed.

[0030] In some embodiments of this application, the intelligent teaching method can be implemented using a processor in an intelligent teaching system. The processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), controller, microcontroller, and microprocessor.

[0031] In this embodiment, the intelligent teaching method can be applied to an intelligent teaching system. The intelligent teaching system is suitable for teaching scenarios, such as teaching scenarios for computer algorithm courses. Its main function is to integrate teaching theory with practical learning to improve teaching efficiency and quality.

[0032] Figure 1 A flowchart of an intelligent teaching method provided in an embodiment of this application is shown below. Figure 1 As shown, the process may include:

[0033] Step 100: Obtain a binocular image containing the user's hand, and determine the position of the user's fingertips based on the binocular image.

[0034] For example, the intelligent teaching system may include an image perception module, which may include an image acquisition subunit and a positioning subunit; wherein, the image acquisition subunit is used to acquire a binocular image containing the user's hand, and the positioning subunit is used to determine the position of the user's fingertips based on the binocular image.

[0035] In this embodiment, the image acquisition subunit can be an image sensor, which can be a binocular camera. The installation location of the binocular camera is not limited; for example, it can be installed on an electronic device or at other locations where it can capture the user's hand. The electronic device can be a head-mounted device or other types of electronic devices, without limitation. It should be noted that when the electronic device is a head-mounted device, the binocular camera can be installed at the center with the bridge of the nose as the axis.

[0036] For example, in a teaching scenario, a binocular camera can be used to capture the user's hand, obtaining a binocular image containing the user's hand; wherein the binocular camera includes a left camera and a right camera, and the binocular image includes the left image I. l And right image I r .

[0037] In this embodiment of the application, the image acquisition subunit obtains the left image I. l And right image I r Then, the positioning subunit can be based on the left image I l And right image I r This determines the position of the user's fingertips. Here, fingertip position refers to the spatial coordinate position of the user's fingertips, that is, the three-dimensional coordinate position of the user's fingertips in the world coordinate system.

[0038] In some embodiments, determining the fingertip position of a user's hand based on a stereo image may include: determining the difference between pairwise adjacent pixel values ​​in the left image; constructing a minimum cost function based on each difference in the left image; solving for a pixel value threshold in the minimum cost function to obtain a pixel value threshold; segmenting the left image based on the pixel value threshold to obtain a foreground segmented image containing the user's hand; and registering the foreground segmented image and the right image to obtain the fingertip position of the user's hand.

[0039] In this embodiment of the application, the positioning subunit can locate the left image I. l The leftmost top-left pixel coordinates are used as the positioning reference point (u0, v0), and the coordinates of all pixels to be located are based on the leftmost pixel coordinates of the image I. l The pixels are calculated, while the right image I r As a necessary condition for triangulation, through the left image I l And right image I r Registration completes the depth measurement of specific pixels. It should be noted that since the hand plays an important role in the operation of the entire teaching intelligent system, the localization subunit will constantly detect and locate the hand. The commonly used method for hand detection is semantic segmentation model, but the inference speed and computing power requirements are not suitable for intelligent teaching systems. Therefore, this application embodiment transforms hand segmentation into the task of foreground and background segmentation; the process is described below by example.

[0040] For example, first, the left image I l As the object of analysis, the hand is used as the foreground (Category A), and the rest is used as the background (Category B); further, the left image I... l As a two-dimensional random signal, rationally speaking, when an image is segmented into two classes, each class itself has low variation between its pixels. Since low variation indicates high homogeneity, the left image I can be defined. l The difference between the values ​​of any two adjacent pixels is shown in formula (1):

[0041]

[0042] Here, s i,j and s i,j+1 This represents the values ​​of two adjacent pixels.

[0043] For example, when determining the left image I according to the above formula (1) l After calculating the differences between any two adjacent pixel values ​​in the left image, a minimum cost function can be constructed based on each difference in the left image. Understandably, in an ideal background and foreground, the difference between two pixel clusters should be minimized. Therefore, a minimum cost function can be constructed based on the differences between two pixel clusters, as shown in formula (2).

[0044]

[0045] in:

[0046]

[0047]

[0048] Here, the minimum cost function includes the pixel value threshold T; TV(f|T) and TV(b|T) represent the total change in pixel values ​​of the foreground and background under the pixel value threshold T, respectively. This represents the pixel value in the foreground. This represents the pixel value in the background. f and p b These represent the proportions of foreground and background pixels in the total number of pixels, respectively.

[0049] In this embodiment of the application, after obtaining the minimum cost function, the pixel value threshold T in the minimum cost function can be solved to obtain the pixel value threshold T; the process may include: obtaining the search range interval of the pixel value threshold; and solving the pixel value threshold in the minimum cost function based on the search range interval to obtain the pixel value threshold.

[0050] For example, a search range interval for a pixel value threshold T can be preset; the search range interval includes a minimum value and a maximum value, wherein the minimum value of the search range interval is greater than that of the left image I. l Minimum pixel value I min The maximum value is less than the left image I. l Maximum pixel value I max Here, the specific values ​​for the search range are not limited; for example, the search range can be set to [I]. min +1,I max -1].

[0051] Furthermore, after obtaining the search range, the pixel value threshold T in the minimum cost function can be solved by iterative search to obtain the pixel value threshold T. The core of this solution method is to find the optimal solution by gradually adjusting the pixel value threshold and evaluating its impact on the minimum cost function, and then determine the optimal solution as the final pixel value threshold T.

[0052] For example, after obtaining the pixel value threshold T, the left image can be segmented based on the pixel value threshold T to obtain a foreground segmentation image containing the user's hand; the corresponding process can be: segmenting the left image I... lThe pixel value of each pixel is compared with a pixel value threshold T. Pixels with a pixel value greater than or equal to the pixel value threshold T are classified as foreground, i.e., foreground segmentation image containing the user's hand. Pixels with a pixel value less than the pixel value threshold T are classified as background.

[0053] Understandably, by analyzing the left image I l By segmenting the foreground and background, the user's hand can be distinguished from the rest of the image. Subsequent registration and other operations can be performed only on the foreground segmentation image containing the user's hand, which can improve the recognition accuracy of the hand and reduce the amount of computation.

[0054] In this embodiment of the application, after obtaining the foreground segmentation image, the foreground segmentation image and the right image can be registered to obtain the fingertip position of the user's hand. This process may include: constructing a region to be registered based on the maximum and minimum values ​​of the vertical coordinates of the foreground segmentation image and the number of pixels in the horizontal direction of the foreground segmentation image; registering the region to be registered with the right image to obtain a registration point pair; performing triangulation on the registration point pair to obtain the depth information of the registration point pair; and determining the fingertip position of the user's hand based on the depth information of the registration point pair.

[0055] Here, the extreme value of the ordinate can be either the maximum or the minimum value, depending on the position of the pixel coordinate origin. If the pixel coordinate origin is at the top left corner of the foreground segmentation image, the extreme value of the ordinate is the minimum value; conversely, if the pixel coordinate origin is at the bottom right corner of the foreground segmentation image, the extreme value of the ordinate is the maximum value.

[0056] For example, the extreme values ​​of the ordinate can be used as the pixel coordinates of the fingertip position (u). hand ,v hand The pixel distance R is determined based on the number of pixels in the horizontal direction of the foreground segmentation image; then, the pixel coordinates (u) are used to... hand ,v hand With R as the center, construct a circle with a pixel distance R as the radius. The circle obtained is the area to be registered.

[0057] In this embodiment of the application, the value of the pixel distance R can be set according to the actual situation, and no specific limitation is made here. For example, 1% of the number of pixels in the horizontal direction of the foreground segmentation image can be set as the pixel distance R.

[0058] Furthermore, in pixel coordinates (u hand ,v hand A circle is constructed with the pixel distance R as the center and the circle as the radius. After obtaining the area to be registered, the circle is used as the feature point to register with the right image to obtain a registration point pair. By performing triangulation on the registration point pair, the depth information of the registration point pair can be obtained. Then, based on the depth information of the registration point pair, the position of the user's fingertips can be determined.

[0059] For example, the registration process refers to the image transformation that takes the largest mutual information by traversing through pixels, which can be achieved by formula (3):

[0060]

[0061] Here, p(p,q) is the joint probability density function, representing the probability of pixel pairs with pixel values ​​p and q in the left and right images, respectively. p(p) and p(q) are the marginal probability density functions of pixels with pixel values ​​p and q in the left and right images, respectively. Ω1 and Ω2 represent the range of pixel values ​​in the left and right images, respectively.

[0062] For example, the trigonometric measurement process refers to solving the least squares solution of formula (4) according to the principle of epipolar geometry:

[0063] s1x1=s2x2+t(4)

[0064] Where t is the baseline value, x1 and x2 are the normalized pixel coordinates of the registration point pair, and s1 and s2 are the depth values ​​of the two points in the registration point pair, i.e., z. hand .

[0065] Furthermore, after obtaining the depth information of the registration point pair, spatial mapping is performed based on the intrinsic parameters of the left camera to obtain the fingertip position of the user's hand, i.e., the spatial coordinate position of the user's fingertip, which is represented as (x... hand ,y hand ,z hand ).

[0066] Step 101: Obtain the text region to be detected within the set range of the fingertip position, and determine the target algorithm type corresponding to the text region to be detected.

[0067] In this embodiment of the application, the intelligent teaching system may further include a text recognition module; wherein, the text recognition module is used to obtain the text region to be detected within a set range of the fingertip position, and determine the target algorithm type corresponding to the text region to be detected; here, the value of the set range can be set according to the actual situation, and is not specifically limited here; for example, the set range can be a 10*10 pixel range near the fingertip position.

[0068] For example, the target algorithm type may include one or more preset algorithm types; wherein, the preset algorithm types may include, but are not limited to, bubble sort algorithm, minimum spanning tree and knapsack problem algorithm, etc.

[0069] For example, after obtaining the text region to be detected, the size of the text region to be detected can be fixed; it should be noted that fixing the size will not have much impact on the recognition result, because the text is relatively short during the lecture.

[0070] In some embodiments, determining the target algorithm type corresponding to the text to be detected may include: acquiring multiple template images; inputting the text region to be detected and the multiple template images into an encoder for template matching to obtain the target template image with the highest similarity; inputting the target template image into a converter to obtain a set of scalar values ​​corresponding to the target template image; and determining the target algorithm type corresponding to the text region to be detected based on the set of scalar values; wherein the set of scalar values ​​includes one or more scalar values.

[0071] For example, the text recognition module obtains the pixel coordinates (u) of the fingertip position. hand ,v hand Using erosion expansion technology and existing edge detection algorithms, the fingertip position (u) is analyzed. hand ,v hand The text region within the set range is detected to obtain the text region range. Here, the text region range is used to generate the template image. The process of generating the template image is illustrated below.

[0072] Understandably, since the carrier of text may be a display document or blackboard used by a teacher, it is difficult to correctly identify the text in a text region under different morphological features. The embodiments of this application reconstruct text recognition into a sequence labeling problem.

[0073] For example, the text region range is taken as a sequence and input into the encoder encoder(·) to obtain the feature sequence ψ = encoder(range). Here, the encoder is composed of Inception modules, which are convolutional neural network structures in deep learning.

[0074] Next, the converter transfer(·,·) transforms the feature sequence ψ into the corresponding set of scalar values ​​ξ = transfer(ψ,s). This set of scalar values ​​includes a set of scalar values, each of which corresponds to a character region. The converter is defined as shown in formula (5):

[0075]

[0076] Among them, l i,s For the features of each character region in the label sequence s, l i,ψ For each character region of the feature sequence ψ, the label sequence s and the feature sequence ψ can be transformed into each other through multiple mappings; ε is a coefficient, and f represents the representation of the feature sequence ψ. For conditional output εl i,ψ , This means taking ε such that l i,s and εl i,ψ The similarity is the highest.

[0077] Furthermore, a posterior distribution P(s|range) containing the feature sequence ψ and the set of scalar values ​​ξ is constructed, as shown in Equation (6):

[0078]

[0079] Where s represents the correct label sequence, s' represents a random combination of label sequences, and E represents the cross-entropy loss function; a total loss function is constructed based on the posterior distribution P(s|range) for model training, and the total loss function is shown in formula (7):

[0080] Loss = 1 - P(s|range) (7)

[0081] Next, the feature coefficients ε in the feature sequence ψ are used to construct a coefficient mapping relationship, resulting in a set of scalar values. The matching set is then used to perform similarity matching on each scalar value in the set, yielding a corresponding matching template. The matching set includes multiple different types of matching templates. Through these steps, various matching templates, i.e., template images, can be obtained.

[0082] In this embodiment, after obtaining multiple template images, the text region to be detected and the multiple template images can be input into the encoder (·) for template matching to obtain the target template image with the highest similarity. Here, the target template image is one of the multiple template images. Then, the target template image is input into the converter (·,·) to obtain the scalar value set ξ corresponding to the target template image. Then, the target algorithm type corresponding to the text region to be detected is determined based on the scalar value set ξ.

[0083] In some embodiments, determining the target algorithm type corresponding to the text region to be detected based on the set of scalar values ​​may include: processing each scalar value in the set of scalar values ​​based on the trained classification model to obtain a prediction result for each scalar value; the prediction result includes prediction probability values ​​for multiple preset algorithm types; and determining the target algorithm type corresponding to the text region to be detected based on the obtained prediction results; the target algorithm type includes at least one preset algorithm type among multiple preset algorithm types.

[0084] For example, the classification model may include a deep learning model for classification and a classification layer; wherein the classification layer is connected to the output of the deep learning model; here, the type of deep learning model is not limited, for example, it may be a Long Short Term Memory (LSTM) network, or other types of deep learning models; the classification layer may be implemented by a softmax function, which is used to map the output of the deep learning model for multiple preset algorithm types to the interval [0 1] to obtain the prediction probability values ​​of multiple preset algorithm types.

[0085] For example, based on the trained classification model, each scalar value in the set of scalar values ​​ξ corresponding to the target template image is processed to obtain a prediction result for each scalar value; that is, each scalar value corresponds to a prediction result; wherein, the prediction result may include prediction probability values ​​of multiple preset algorithm types; here, multiple preset algorithm types are various algorithm types that the model can recognize in advance; the preset algorithm types may include, but are not limited to, bubble sort algorithm, minimum spanning tree and knapsack problem algorithm, etc.

[0086] For example, after obtaining the prediction result for each scalar value, the maximum value among multiple predicted probability values ​​is determined, and the preset algorithm type corresponding to the maximum value is determined as the algorithm type corresponding to the scalar value; the set of algorithm types corresponding to each scalar value is determined as the target algorithm type corresponding to the text region to be detected; that is, the target algorithm type includes the algorithm type corresponding to each scalar value.

[0087] For example, suppose the prediction results for a certain scalar value are: the predicted probability values ​​for bubble sort algorithm, minimum spanning tree algorithm, and knapsack problem algorithm are 0.2, 0.3, and 0.5 respectively. Since the predicted probability value of knapsack problem algorithm is 0.5, which is the maximum of the three predicted probability values, knapsack problem algorithm can be identified as the algorithm type corresponding to this scalar value. Similarly, the algorithm types corresponding to other scalar values ​​can be obtained. Then, based on the algorithm types corresponding to these scalar values, the target algorithm type corresponding to the text region to be detected can be determined.

[0088] Step 102: Extract each logic block related to the target algorithm type from the logic block library. When a parameter setting trigger operation is detected for the target logic block in each logic block, obtain the target logic block after parameter setting.

[0089] In this embodiment of the application, the logic block library may include multiple encapsulated logic blocks for implementing different algorithm functions; the logic block library may be pre-built or built when this step is performed; for example, the logic block library may include, but is not limited to, "loop block", "judgment block", "operation block", "input block" and "output block".

[0090] For example, the intelligent teaching system may also include a logic construction module; after obtaining the target algorithm type corresponding to the text region to be detected according to the above steps, the logic construction module can extract each logic block related to the target algorithm type from the logic block library; as can be seen from the above, the target algorithm type may include at least one preset algorithm type. Therefore, after obtaining the target algorithm type corresponding to the text region to be detected, the various preset algorithm types included in the target algorithm type can be used as indexes to extract each logic block related to it from the logic block library.

[0091] For example, if the target algorithm type includes the bubble sort algorithm, the logic blocks extracted from the logic block library include "loop block", "judgment block", "input block" and "output block"; if the target algorithm type includes the minimum spanning tree, the logic blocks extracted from the logic block library include "loop block", "operation block", "input block" and "output block"; if the target algorithm type includes the knapsack problem algorithm, the logic blocks extracted from the logic block library include "loop block", "judgment block", "operation block", "input block" and "output block".

[0092] For example, the extracted logic blocks can be displayed as squares composed of different colors and character identifiers. For instance, the squares corresponding to each logic block can be displayed on the left side of the perceptual field of view (the area where the left image is located) through a visualization module. This process can be achieved through holographic projection.

[0093] In some embodiments, detecting a parameter setting trigger operation for a target logic block in each logic block may include: determining the number of consecutive changes in depth information; and determining that a parameter setting trigger operation for a target logic block in each logic block has been detected when the number of consecutive changes reaches a set number.

[0094] Here, there is no specific limitation on the value of the set number of times. For example, the set number of times can be two or three times.

[0095] The parameter setting trigger operation can be understood as a trigger operation used to set the logic parameters of the target logic block. The triggering method corresponding to this trigger operation can be set according to the set number of times. For example, if the set number of times is three, the parameter setting trigger operation is a three-click operation, that is, it can be triggered by a three-click operation.

[0096] For example, as can be seen from the above, the position of a user's fingertips includes depth information, namely, the depth value (z). hand ); understandably, with each click of the user's fingertip, the depth value (z) handThe number of consecutive changes in depth information, even a single change, can represent the number of consecutive clicks by the user's fingertip. Therefore, the number of changes in the depth value can be used to determine whether a user-triggered operation targeting the parameter settings of a target logic block within each logic block has been detected.

[0097] In this embodiment, the number of consecutive changes in depth information can be determined and compared with a set number. If the number of consecutive changes reaches the set number, it is determined that a user has triggered a parameter setting operation for the target logic block in each logic block. In this case, parameter adjustment for the target logic block in each logic block is triggered. Conversely, if the number of consecutive changes does not reach the set number, it is determined that no user has triggered a parameter setting operation for the target logic block in each logic block. In this case, parameter adjustment for the target logic block in each logic block is not triggered.

[0098] For example, suppose the number of clicks is set to three, that is, the parameter setting triggers a three-click operation; if a depth value (z) is detected... hand If the parameter settings change three times consecutively, it is determined that a user-triggered operation for setting parameters of the target logic block in each logic block has been detected; otherwise, it is determined that no user-triggered operation for setting parameters of the target logic block in each logic block has been detected.

[0099] In this embodiment of the application, when the logic construction module detects that the user has triggered a parameter setting operation for the target logic block in each logic block, it can notify the visualization module to pop up a virtual keyboard and a parameter setting window for the target logic block. The user can use the virtual keyboard to set the relevant parameters of the target logic block. At this time, the target logic block after parameter setting can be obtained.

[0100] For example, when the target logic block is a "loop block", the user can use the virtual keyboard to set the number of loops and the loop variable in the parameter setting window; at this time, the "loop block" after parameter setting can be obtained.

[0101] As can be seen, in this embodiment of the application, users can set the parameters of the logic block themselves through the pop-up parameter setting window, which can meet their own learning needs.

[0102] In some embodiments, before the detection parameter setting trigger operation, the above method may further include: when a movement trigger operation is detected for a target logic block in each logic block, determining the current fingertip position of the user's hand; determining whether the current fingertip position overlaps with the pixel of the target logic block, and if so, generating a copy logic block of the target logic block; determining the position change of the current fingertip position, and controlling the copy logic block to perform the corresponding target operation according to the position change.

[0103] The move trigger operation can be understood as a trigger operation used to adjust the position of the target logic block. In the specific implementation, the triggering method corresponding to this trigger operation can be set according to the actual situation; for example, it can be triggered by a sliding operation.

[0104] For example, the target logic block is a logic block that requires parameter settings, and there can be one or more target logic blocks; when a user triggers a target logic block, the current fingertip position (x) of the user's hand can be determined first. hand ,y hand ,z hand Then, based on the current fingertip position (x) hand ,y hand Determine whether the current fingertip position overlaps with the pixel of the target logic block. If so, generate a copy logic block of the target logic block.

[0105] Furthermore, the positional changes of the current fingertip can be determined, and then the copy logic block can be controlled to perform corresponding target operations based on the positional changes; here, the target operation may include at least one of movement and scaling. For example, the above process can be: determining the current fingertip position (x... hand ,y hand After a change occurs, guided by overlapping pixels, all pixels in the resulting copy logic block follow (x) hand ,y hand The movement is based on changes in the position of the user's fingertip. For example, when the position of the user's current fingertip changes (x... hand ,y hand If the fingertip moves to position (1, 1), the overlapping pixels of the copied logical block also move to position (1, 1), and the remaining pixels follow. The depth value (z) of the current fingertip position is then determined. hand When the value of the copy logic block changes, the corresponding scaling is controlled. Here, the scaling pattern of the copy logic block is not limited; for example, it can be adjusted based on the depth value (z). hand When the value (z) decreases, the control block for replication logic decreases, and the depth value (z) becomes smaller. hand When the value increases, the control copy logic block increases in size.

[0106] As can be seen, the embodiments of this application can visualize the extracted logic blocks, and users can adaptively adjust the position of the logic blocks according to their own needs, which can meet their own learning needs.

[0107] Step 103: Merge the target logic block after parameter setting and the logic block without parameter setting, obtain the fusion result and display the fusion result.

[0108] In this embodiment of the application, the intelligent teaching system may further include a visualization module; after obtaining the target logic block with parameters set according to the above steps, the target logic block with parameters set and the logic block without parameters set can be fused to obtain a fusion result, and the fusion result can be displayed through the visualization module.

[0109] Here, the logic blocks without parameter settings include all extracted logic blocks except the target logic block. For example, fusion refers to stacking the target logic block and the logic blocks without parameter settings together according to the algorithm logic to complete the final algorithm function.

[0110] For example, assuming the target algorithm is a knapsack problem algorithm, the following fusion result can be obtained by fusing the various logic blocks related to it: "loop block", "judgment block", "operation block", "input block" and "output block"; here, the various logic blocks related to the knapsack problem algorithm include the target logic block after parameter settings.

[0111]

[0112]

[0113] In this embodiment of the application, after obtaining the fusion result, the fusion result can be displayed through the visualization module; for example, the visualization module can complete the projection of relevant content through holographic projection, and all projected content can be transformed according to the logic construction module; at the same time, the image perception module will synchronously acquire the content projected by the visualization module.

[0114] As can be seen, the intelligent teaching method of this application first determines the position of the user's fingertip, then determines the target algorithm type corresponding to the text region to be detected pointed to by the fingertip, and then extracts each logic block related to the target algorithm type. The user sets the parameters of each logic block to realize the corresponding algorithm function. The whole process can be carried out while the user learns the theoretical logic. In this way, it can ensure that the user completes the practical learning of the relevant algorithm while mastering the theoretical logic. That is, it can realize the integration between teaching theory and practical learning, which is conducive to improving teaching efficiency and quality and ensuring teaching effect.

[0115] Figure 2 This is a schematic diagram of the composition structure of the intelligent teaching system according to an embodiment of this application, such as... Figure 2 As shown, the system includes: an image perception module 200, a text recognition module 201, a logic construction module 202, and a visualization module 203, wherein:

[0116] The image perception module 200 is used to acquire a binocular image containing the user's hand, and determine the position of the user's fingertips based on the binocular image;

[0117] The text recognition module 201 is used to acquire the text region to be detected within a set range of the fingertip position and determine the target algorithm type corresponding to the text region to be detected;

[0118] The logic construction module 202 is used to extract various logic blocks related to the target algorithm type from the logic block library, and when a parameter setting trigger operation is detected for the target logic block in each logic block, the target logic block after parameter setting is obtained; the target logic block after parameter setting and the logic block without parameter setting are fused to obtain the fusion result.

[0119] Visualization module 203 is used to display the fusion results.

[0120] In some embodiments, the text recognition module 201 is further configured to:

[0121] Obtain multiple template images;

[0122] The text region to be detected and multiple template images are input into the encoder for template matching to obtain the target template image with the highest similarity; the target template image is one of the multiple template images.

[0123] The target template image is input into the converter to obtain a set of scalar values ​​corresponding to the target template image; the set of scalar values ​​includes one or more scalar values.

[0124] The target algorithm type corresponding to the text region to be detected is determined based on the set of scalar values.

[0125] In some embodiments, the text recognition module 201 is further configured to:

[0126] The trained classification model processes each scalar value in the scalar value set to obtain a prediction result for each scalar value; the prediction result includes prediction probability values ​​for various preset algorithm types.

[0127] Based on the obtained prediction results, the target algorithm type corresponding to the text region to be detected is determined; the target algorithm type includes at least one of a variety of preset algorithm types.

[0128] In some embodiments, before the detection parameter setting trigger operation, the logic construction module 202 is further configured to:

[0129] When a movement trigger operation is detected targeting a target logic block in each logic block, the current fingertip position of the user's hand is determined;

[0130] Determine whether the current fingertip position overlaps with the pixel of the target logic block. If so, generate a copy logic block of the target logic block.

[0131] Determine the position change of the current fingertip, and control the copy logic block to execute the corresponding target operation based on the position change; the target operation includes at least one of movement and scaling.

[0132] In some embodiments, the binocular image includes a left image and a right image, and the image perception module 200 is further configured to:

[0133] Determine the difference between any two adjacent pixel values ​​in the left image;

[0134] Based on the differences in the left image, a minimum cost function is constructed; the minimum cost function includes pixel value thresholds.

[0135] The pixel value threshold is obtained by solving the minimum cost function;

[0136] The left image is segmented based on a pixel value threshold to obtain a foreground segmentation image containing the user's hand;

[0137] The foreground segmentation image and the right image are registered to obtain the position of the user's fingertips.

[0138] In some embodiments, the image sensing module 200 is further configured to:

[0139] Obtain the search range for the pixel value threshold; the minimum value of the search range is greater than the minimum pixel value of the left image, and the maximum value is less than the maximum pixel value of the left image.

[0140] Based on the search range, the pixel value threshold in the minimum cost function is solved to obtain the pixel value threshold.

[0141] In some embodiments, the image sensing module 200 is further configured to:

[0142] Based on the maximum and minimum values ​​of the ordinate of the foreground segmentation image and the number of pixels in the horizontal direction of the foreground segmentation image, the region to be registered is constructed.

[0143] Register the region to be registered with the right image to obtain a registration point pair;

[0144] Triangulation is performed on the registration point pair to obtain the depth information of the registration point pair;

[0145] The position of the user's fingertips is determined based on the depth information of the registration point pair.

[0146] In some embodiments, the fingertip position of the user's hand includes depth information, and the logic construction module 202 is further configured to:

[0147] Determine the number of consecutive changes in depth information;

[0148] When the number of consecutive changes reaches a set number, it is determined that a parameter setting trigger operation for the target logic block in each logic block has been detected.

[0149] In practical applications, the image perception module 200, text recognition module 201, logic construction module 202 and visualization module 203 can all be implemented by a processor located in an electronic device. The processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller and microprocessor.

[0150] Furthermore, in this embodiment, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.

[0151] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0152] Specifically, the computer program instructions corresponding to an intelligent teaching method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the computer program instructions corresponding to an intelligent teaching method in the storage media are read or executed by an electronic device, any one of the intelligent teaching methods in the aforementioned embodiments is implemented.

[0153] Based on the same technical concept as the foregoing embodiments, see Figure 3 It illustrates an electronic device 300 provided in an embodiment of this application, which may include: a memory 301 and a processor 302; wherein,

[0154] Memory 301 is used to store computer programs and data;

[0155] The processor 302 is configured to execute a computer program stored in the memory to implement any of the intelligent teaching methods described in the foregoing embodiments.

[0156] In practical applications, the memory 301 mentioned above can be volatile memory, such as RAM; or non-volatile memory, such as ROM, flash memory, hard disk drive (HDD) or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 302.

[0157] The processor 302 described above can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor. It is understood that for different intelligent teaching systems, the electronic device used to implement the functions of the processor described above can also be other types, and this application embodiment does not specifically limit the specific implementation.

[0158] This application also provides a computer storage medium storing a computer program; when the computer program is executed, it can implement any of the intelligent teaching methods described in the foregoing embodiments.

[0159] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the intelligent teaching methods described in the foregoing embodiments.

[0160] In some embodiments, the system provided in this application may have functions or include modules that can be used to execute the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0161] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0162] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.

[0163] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0164] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0165] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0166] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.

[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0168] The above are merely preferred embodiments of this application and are not intended to limit the scope of protection of this application.

Claims

1. An intelligent teaching method, characterized in that, The method, applied to an intelligent teaching system, includes: Acquire a binocular image containing the user's hand, and determine the fingertip position of the user's hand based on the binocular image; Obtain the text region to be detected within the set range of the fingertip position, and determine the target algorithm type corresponding to the text region to be detected; Extract each logic block related to the target algorithm type from the logic block library, and when a parameter setting trigger operation is detected for the target logic block in each logic block, obtain the target logic block after parameter setting. The target logic block with parameters set and the logic block without parameters set are fused together to obtain a fusion result, which is then displayed.

2. The method according to claim 1, characterized in that, Determining the target algorithm type corresponding to the text region to be detected includes: Obtain multiple template images; The text region to be detected and the various template images are input into the encoder for template matching to obtain the target template image with the highest similarity; the target template image is one of the various template images. The target template image is input into the converter to obtain a set of scalar values ​​corresponding to the target template image; the set of scalar values ​​includes one or more scalar values. The target algorithm type corresponding to the text region to be detected is determined based on the set of scalar values.

3. The method according to claim 2, characterized in that, The step of determining the target algorithm type corresponding to the text region to be detected based on the scalar value set includes: The training-complete classification model processes each scalar value in the scalar value set to obtain a prediction result for each scalar value; the prediction result includes prediction probability values ​​of various preset algorithm types; Based on the obtained prediction results, the target algorithm type corresponding to the text region to be detected is determined; the target algorithm type includes at least one preset algorithm type among the multiple preset algorithm types.

4. The method according to claim 1, characterized in that, Before detecting the parameter setting trigger operation, the method further includes: When a movement trigger operation is detected targeting a target logic block in each of the logic blocks, the current fingertip position of the user's hand is determined; Determine whether the current fingertip position overlaps with the pixel of the target logic block. If so, generate a copy logic block of the target logic block. Determine the positional change of the current fingertip position, and control the copy logic block to perform the corresponding target operation based on the positional change; the target operation includes at least one of movement and scaling.

5. The method according to claim 1, characterized in that, The binocular images include a left image and a right image. Determining the fingertip position of the user's hand based on the binocular images includes: Determine the difference between any two adjacent pixel values ​​in the left image; Based on the differences in the left image, a minimum cost function is constructed; the minimum cost function includes pixel value thresholds. The pixel value threshold is obtained by solving the minimum cost function; The left image is segmented based on the pixel value threshold to obtain a foreground segmentation image containing the user's hand; The foreground segmentation image and the right image are registered to obtain the fingertip position of the user's hand.

6. The method according to claim 5, characterized in that, Solving for the pixel value threshold in the minimum cost function to obtain the pixel value threshold includes: Obtain the search range interval of the pixel value threshold; the minimum value of the search range interval is greater than the minimum pixel value of the left image, and the maximum value is less than the maximum pixel value of the left image; Based on the search range, the pixel value threshold in the minimum cost function is solved to obtain the pixel value threshold.

7. The method according to claim 5, characterized in that, The step of registering the foreground segmentation image and the right image to obtain the fingertip position of the user's hand includes: Based on the maximum and minimum values ​​of the vertical coordinates of the foreground segmentation image and the number of pixels in the horizontal direction of the foreground segmentation image, a region to be registered is constructed. The region to be registered and the right image are registered to obtain a registration point pair; Triangulation is performed on the registration point pair to obtain the depth information of the registration point pair; The position of the user's fingertips is determined based on the depth information of the registration point pair.

8. The method according to claim 1, characterized in that, The fingertip position of the user's hand includes depth information, and the detection of parameter setting trigger operations for target logic blocks in each logic block includes: Determine the number of consecutive changes in the depth information; When the number of consecutive changes reaches a set number, it is determined that a parameter setting trigger operation for the target logic block in each logic block has been detected.

9. An intelligent teaching system, characterized in that, The system includes: An image perception module is used to acquire a binocular image containing the user's hand, and to determine the position of the fingertips of the user's hand based on the binocular image; The text recognition module is used to acquire the text region to be detected within a set range of the fingertip position, and to determine the target algorithm type corresponding to the text region to be detected; The logic construction module is used to extract various logic blocks related to the target algorithm type from the logic block library, and when a parameter setting trigger operation is detected for the target logic block in each logic block, the target logic block after parameter setting is obtained; the target logic block after parameter setting and the logic block without parameter setting are fused to obtain a fusion result. The visualization module is used to display the fusion results.

10. An electronic device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method according to any one of claims 1 to 8.

11. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method described in any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.