Voice reading method and system for multi-modal mathematical formula recognition and medium

By employing a multimodal input method for reading mathematical formulas by speech, and utilizing a visual recognition module to extract features from images, handwritten trajectories, and geometric figures, combined with a mathematical corpus for semantic parsing, the problem of incomplete multimodal input and low parsing accuracy is solved, achieving comprehensive and accurate speech conversion of mathematical formulas.

CN121808348APending Publication Date: 2026-04-07JIANGSU HAOHAN INFORMATION TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing mathematical formula voice reading technologies suffer from incomplete multimodal input coverage and low semantic parsing accuracy, making it unable to effectively interpret the structural meaning of geometric formulas and handwritten formulas with non-standard typesetting or writing deviations.

Method used

By acquiring multimodal inputs such as images, handwritten trajectories, and geometric figures, the system uses a visual recognition module to extract formula text features and morphological features, and combines these with a mathematical corpus for semantic parsing and speech synthesis to generate natural language text.

Benefits of technology

It improves the comprehensiveness of multimodal input coverage and the accuracy of formula semantic parsing, and can accurately identify and convert the speech output of various mathematical formulas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808348A_ABST
    Figure CN121808348A_ABST
Patent Text Reader

Abstract

The invention discloses a voice reading method and system for multi-mode mathematical formula recognition and a medium, and relates to the related field of data processing, and the method comprises the steps: obtaining mathematical formula input modes, including images, handwriting tracks and geometric figures; performing feature extraction on the input multi-modal information through a visual identification module, and identifying formula character features and formula morphological features; the formula character features and the formula morphological features are used for analysis, and formula structure semantic information is recognized; and carrying out natural language voice conversion based on the formula structure semantic information in combination with a mathematical field corpus, and reading voice information. The technical problems of incomplete multi-modal input coverage and low semantic analysis precision in the existing mathematical formula voice reading technology are solved, and the technical effect of improving multi-modal input coverage comprehensiveness and formula semantic analysis precision is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to speech reading methods, systems and media for multimodal mathematical formula recognition. Background Technology

[0002] In industrial and public sectors such as digital education, intelligent document processing, and accessible information services, speech recognition technology for mathematical formulas is a crucial support for bridging the gap between visual and auditory information. Its maturity directly impacts the service capabilities of products like intelligent teaching systems and document voice assistants, demonstrating significant industrial and social value. Currently, mainstream technologies for speech recognition of mathematical formulas primarily focus on processing single-modal inputs. Specifically, this involves extracting characters from printed formula images using optical character recognition (OCR) or parsing symbols from regular handwritten formulas using handwriting trajectory recognition algorithms. The resulting character sequences are then matched to a fixed speech synthesis template to generate the final speech output. However, existing methods do not fully cover diverse and heterogeneous formula input types, and their feature extraction and semantic parsing logic is relatively simplistic, focusing only on the character symbol information of the formula while neglecting its spatial form and structural relationships. This results in an inability to effectively interpret the structural meaning of geometric formulas or accurately recognize the semantics of non-standard typesetting formulas or handwritten formulas with writing errors, ultimately leading to speech conversion results that do not match the true mathematical meaning of the formula.

[0003] Currently, mathematical formula speech reading technology suffers from technical problems such as incomplete multimodal input coverage and low semantic parsing accuracy. Summary of the Invention

[0004] This application provides a speech reading method, system, and medium for multimodal mathematical formula recognition. It acquires arbitrary mathematical formula input modalities from images, handwritten trajectories, and geometric figures. A visual recognition module extracts features from the input information of this modality, obtaining both textual and morphological features of the formula. These two types of features are then analyzed to identify the structural semantic information of the formula. Combined with a mathematical corpus, the structural semantic information of the formula is converted into natural language text. The corresponding speech information is then generated through speech synthesis for reading. These techniques solve the technical problems of incomplete multimodal input coverage and low semantic parsing accuracy in existing mathematical formula speech reading technologies, achieving the technical effect of improving the comprehensiveness of multimodal input coverage and the accuracy of formula semantic parsing.

[0005] This application provides a speech reading method for multimodal mathematical formula recognition, comprising: acquiring the mathematical formula input modality, including images, handwritten trajectories, and geometric figures; extracting features from the input multimodal information through a visual recognition module to identify formula text features and formula morphological features; parsing the formula text features and formula morphological features respectively to identify formula structural semantic information; and performing natural language speech conversion based on the formula structural semantic information combined with a mathematical corpus to read the speech information.

[0006] In a possible implementation, the formula text features and formula form features are analyzed separately to identify the formula structure semantic information, and the following processing is performed: semantic analysis is performed on the formula text features to identify the formula structure semantic information; data formula structure form matching analysis is performed on the formula form features, and the mathematical formula obtained from the matching analysis is used to identify the formula structure semantic information.

[0007] In a possible implementation, a visual recognition module extracts features from the input multimodal information, identifies formula text features and formula shape features, and performs the following processing: the image, handwritten trajectory, and geometric figure are converted into an image to obtain image information; the image information is used to recognize mathematical characters, symbols, and numbers through an optical character recognition module to obtain standardized symbols; the geometric figure is used to identify edges and graphic structure, and a mapping relationship between standardized symbols in the image information and graphic structure is established; the standardized symbols are used as formula text features, and the graphic structure features and their mapping relationship are used as formula shape features.

[0008] In a possible implementation, the image, handwritten trajectory, and geometric figure are subjected to image recognition conversion to obtain image information, and the following processing is performed: The input image is preprocessed, including removing image noise, adjusting image brightness and contrast, rotating and cropping the image to ensure standardization; the handwritten trajectory input is smoothed to remove irregular noise and reconstruct the strokes of the trajectory, converting it into a standardized image format; the geometric figure input undergoes edge detection and structure distribution extraction, extracting the geometric boundaries of the figure and converting it into standardized image information, wherein geometric annotation mapping relationships are established for figures with character identifiers; the preprocessed image information, the standardized image information converted from the handwritten trajectory, and the standardized image information converted from the geometric figure are output to obtain the image information.

[0009] In a possible implementation, semantic parsing is performed on the formula text features to identify formula structural semantic information, and the following processing is performed: based on the formula text features, the identified mathematical symbols are classified, and the types of mathematical symbols include at least operators, variables, constants, and numbers; various types of mathematical symbols are converted into standard encoding formats to form standard encoding sequences; based on the standard encoding sequences, the mathematical symbols in the formula are parsed to construct a formula syntax tree; based on the formula syntax tree, the hierarchical structure and operational relationships between mathematical symbol types are identified to generate formula structural semantic information.

[0010] In a possible implementation, the formula morphological features are analyzed using data formula structure morphological matching. The mathematical formulas obtained from the matching analysis are used to identify the semantic information of the formula structure, and the following processing is performed: Based on the graphic structure features, a search and matching process is conducted in a pre-defined mathematical formula template learning space to determine the basic mathematical formula or theorem definition that matches the current graphic; based on the annotation mapping relationship, the annotation characters identified in the graphic are substituted as specific parameters into the matched basic mathematical formula or theorem template to instantiate and generate a target mathematical expression corresponding to the current graphic; the target mathematical expression is subjected to mathematical symbol recognition and grammatical analysis to clarify the internal operational priorities and structural relationships, and the components of the symbolic expression are mapped back to the original graphic structure features to establish a semantic association between symbolic elements and geometric primitives, thereby obtaining the semantic information of the formula structure.

[0011] In a possible implementation, natural language speech conversion is performed based on the semantic information of the formula structure combined with a mathematical corpus. The speech information is read and the following processing is performed: traversing the semantic information of the formula structure, and linearizing the hierarchical semantic structure into a primary natural language text sequence according to preset mathematical reading rules. This primary natural language text sequence is a structured natural language that conforms to spoken language habits. The mathematical corpus is invoked to perform semantic parsing and replacement optimization on the mathematical objects in the primary natural language text sequence, resulting in a natural language description text. This includes: applying standard reading rules to basic mathematical symbols and function names; matching and applying common expressions or theorem names to formula patterns or geometric relationships; adding descriptive qualifiers to variables and constants based on contextual information to make the description more complete and unambiguous; automatically inserting corresponding prosodic markers, including pauses, emphasis, and speech rate adjustment, based on the inherent mathematical logic structure of the natural language description text; and inputting the prosodic-marked natural language description text into a speech synthesis engine to generate speech information.

[0012] In a possible implementation, the following processing is also performed: matching the input formula modality with a preset modality to determine the modality matching relationship; when no modality matching relationship exists, identifying the correlation and difference features between the current input formula modality and the preset modality; based on the correlation features, mapping the data of the input modality to the feature extraction and recognition process corresponding to the preset modality with the highest correlation; based on the difference features, performing data transformation or compensation on the input modality through an intermediate conversion module to generate an intermediate representation that meets the processing requirements of the preset modality; and using the mapped preset modality recognition process, performing mathematical formula recognition processing on the intermediate representation to obtain mathematical semantic information.

[0013] This application also provides a speech reading system for multimodal mathematical formula recognition, comprising: a mathematical formula input modality acquisition module for acquiring mathematical formula input modalities, including images, handwritten trajectories, and geometric figures; a visual recognition module for extracting features from the input multimodal information and recognizing formula text features and formula morphological features; a formula structure semantic information recognition module for parsing the formula text features and formula morphological features respectively and recognizing formula structure semantic information; and a natural language speech conversion module for performing natural language speech conversion based on the formula structure semantic information combined with a mathematical corpus to read speech information.

[0014] This application also provides a computer-readable storage medium, including: a computer program stored thereon, which, when executed by a processor, implements a speech reading method for recognizing multimodal mathematical formulas.

[0015] The proposed method, system, and medium for multimodal mathematical formula recognition speech reading first acquire the mathematical formula input modality, including images, handwritten trajectories, and geometric figures. Then, a visual recognition module extracts features from the input multimodal information, recognizing the formula's textual and morphological features. These features are then analyzed to identify the formula's structural semantic information. Finally, based on this structural semantic information and a mathematical corpus, natural language speech conversion is performed to read the speech information. Through this process, the proposed method, system, and medium achieve the technical effects of improving the comprehensiveness of multimodal input coverage and the accuracy of formula semantic parsing. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0017] Figure 1 This is a schematic flowchart of a speech reading method for multimodal mathematical formula recognition provided in an embodiment of this application.

[0018] Figure 2 A schematic diagram of the structure of a speech reading system for multimodal mathematical formula recognition provided in an embodiment of this application.

[0019] Figure labeling: Mathematical formula input modality acquisition module 10, visual recognition module 20, formula structure semantic information recognition module 30, natural language speech conversion module 40. Detailed Implementation

[0020] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0021] This application provides a speech reading method for multimodal mathematical formula recognition, such as... Figure 1 As shown, the method includes: Step S100: Obtain mathematical formula input modalities, including images, handwritten trajectories, and geometric figures.

[0022] Specifically, the system receives multimodal data from external input interfaces, including image acquisition tools, handwritten trajectory acquisition sensor driver tools, and geometric vector data receiving tools. The input data undergoes format validation to ensure it falls within the preset categories of image, handwritten trajectory, or geometric figure. Specifically, image data requires validation of resolution and color mode, handwritten trajectory data requires validation of the coordinate sequence format of trajectory points, and geometric figure data requires validation of the format of structural information such as vertices, edges, and faces.

[0023] For example, in the application scenario of voice reading of mathematical formulas on educational tablet devices, image modal input can capture the formulas on the paper math test paper through the device's camera and obtain an RGB color mode image; handwritten trajectory modal input can collect the trajectory of a quadratic equation handwritten by the user through the device's capacitive touch screen and obtain a trajectory sequence composed of continuous coordinate points; geometric shape modal input can receive the right triangle vector data drawn by the user through the device's built-in geometric drawing software, including the coordinates of the three vertices, the lengths of the three sides, and right angle identification information.

[0024] Step S200: The visual recognition module extracts features from the input multimodal information and identifies formula text features and formula shape features.

[0025] Specifically, the visual recognition module selects the corresponding feature extraction branch for specific modal data input, such as images, handwritten trajectories, or geometric figures. First, it converts the input data into a unified feature input format through a modality adaptation layer. Then, it extracts low-level visual features through a convolutional neural network. Finally, it outputs the formula text features and formula shape features through a dual-output classifier. The model used is a convolutional neural network-based feature extraction model, which includes an input layer, a modality adaptation layer, convolutional layers, pooling layers, and fully connected layers.

[0026] For example, in the application scenario of voice reading of mathematical formulas on educational tablet devices, if the input is an image modality, such as an image of a quadratic equation on a paper test paper, the visual recognition module activates the image feature extraction branch, and the modality adaptation layer directly receives the image data; if the input is a handwritten trajectory modality, such as the trajectory of a Pythagorean theorem handwritten by the user, the modality adaptation layer converts the trajectory coordinate sequence into a binary image format; if the input is a geometric shape modality, such as circular vector data drawn by the user, including the center coordinates and radius, the modality adaptation layer converts the vector data into a contour image format. In all three cases, the convolutional layer extracts edge and texture features from the processed input, the pooling layer reduces dimensionality, and the fully connected layer finally outputs formula text features and formula shape features, such as unknowns and coefficients in the equation, and annotation characters in the graph; and formula shape features such as the writing structure of the equation and the contour shape of the circle.

[0027] In one possible implementation, a visual recognition module extracts features from the input multimodal information, recognizing formula text features and formula shape features. Step S200 further includes step S210, which performs image conversion on the image, handwritten trajectory, and geometric figure to obtain image information. Specifically, a corresponding conversion algorithm is selected for the input modality type. Image modalities are directly preprocessed with standardization, handwritten trajectory modalities are converted into images through trajectory reconstruction, and geometric figure modalities are converted into images through contour extraction. Finally, all conversion results are unified into a grayscale image format.

[0028] For example, in the scenario of recognizing elementary school math formulas, if the input is a fuzzy image, Gaussian filtering is directly applied to the image for noise reduction and brightness adjustment; if the input is a handwritten trajectory, the trajectory is smoothed using the moving average method, and the strokes are reconstructed using Bézier curves and converted into a binary image; if the input is a geometric shape, the square contour is extracted using Canny edge detection and converted into a contour image.

[0029] Step S220: The optical character recognition module performs mathematical character, symbol, and number recognition on the image information to obtain standardized symbols. Specifically, the optical character recognition module adopts a deep learning architecture and sequentially performs five steps: image binarization, tilt correction, feature extraction, character classification, and standardized output. First, the input image is binarized and tilt corrected. Then, character features are extracted through a residual network. Next, character classification is performed through a fully connected layer. Finally, the classification result is converted into a standardized mathematical symbol encoding based on Unicode, such as x corresponding to the encoding U+0078 and square corresponding to the encoding U+00B.

[0030] Step S230 involves edge and structure recognition of the geometric figure, and establishing a mapping relationship between standardized symbols in the image information and the annotations of the geometric structure. Specifically, when the input is a geometric figure modality, Hough transform is used to detect the edges of the figure, such as lines, circles, and triangles. The structural parameters of the figure, such as vertex coordinates, side lengths, and center positions, are determined by least squares fitting. Then, template matching is used to match the standardized symbols with the annotation positions in the geometric structure, establishing a one-to-one mapping relationship.

[0031] For example, in a middle school geometric shape recognition scenario, if the input to step S210 is a geometric shape modality, such as right-angled triangle vector data drawn by the user, labeled A, B, and C, with the right angle located at point C, the positions of the three sides and the right angle of the triangle are first detected using Hough transform to determine that the shape is a right-angled triangle. Then, the pixel positions of the labeled characters A, B, and C are determined using least squares fitting. Finally, a label mapping relationship is established between the standardized symbol A and the triangle vertex A, the standardized symbol B and the vertex B, and the standardized symbol C and the right-angle vertex C. If the input to step S210 is an image or handwritten trajectory modality, and the image does not contain geometric elements, such as a pure formula text image, then this step only outputs the verification result without geometric structure and does not establish a mapping relationship.

[0032] Step S240: The standardized symbols are used as formula text features, and the graphic structure features and their annotation mapping relationships are used as formula morphological features. Specifically, feature encapsulation is used to classify and encapsulate the standardized symbols, graphic structure features, and annotation mapping relationships. The standardized symbols are organized into a formula text feature dataset according to character type, in JSON format. If the input is a geometric modality, the graphic structure features and annotation mapping relationships are integrated into a formula morphological feature dataset. If the input is an image or handwritten trajectory modality, the formula morphological features only include the spatial distribution structure of the text, such as the row arrangement of the formula and the size ratio of the characters.

[0033] In one possible implementation, the image, handwritten trajectory, and geometric figure are transformed into a recognition image to obtain image information. Step S210 further includes step S211, which preprocesses the input image. The preprocessing includes removing image noise, adjusting image brightness and contrast, rotating and cropping the image to ensure image standardization. Specifically, the input image modal data is processed sequentially in the order of noise removal, brightness and contrast adjustment, and rotation and cropping. Noise removal uses a median filtering algorithm, brightness and contrast adjustment uses a gamma correction algorithm, and rotation and cropping uses Hough transform to detect the image tilt angle and perform rotation. The cropping area is determined by edge detection, and finally, a standardized grayscale image is output.

[0034] Step S212 involves smoothing the handwritten trajectory input to remove irregular noise and reconstructing the strokes of the trajectory, converting it into a standardized image format. Specifically, the smoothing process uses a moving average method to remove jitter noise from the handwritten trajectory, the stroke reconstruction uses a Bezier curve fitting algorithm to connect discrete trajectory points into continuous strokes, and finally, a rasterization algorithm converts the vector strokes into a binary image format, with the stroke width set to two pixels, the background set to black, and the strokes set to white.

[0035] Step S213 involves edge detection and structural distribution extraction of the geometric input, extracting the geometric boundaries of the graphic and converting them into standardized image information. For graphics with character identifiers, a geometric annotation mapping relationship is established. Specifically, Canny edge detection is used to extract the boundary contours of the geometric graphics, and contour tracking is used to determine the structural distribution information of the graphics, such as vertices, edges, and faces. For graphics with character identifiers, template matching is used to match the characters with the graphic's annotation positions to establish a geometric annotation mapping relationship. Finally, a rasterization algorithm is used to convert the geometric boundaries into grayscale image information.

[0036] Step S214: Output the preprocessed image information, the standardized image information converted from handwritten trajectories, and the standardized image information converted from geometric shapes to obtain the image information. Specifically, the standardized image information is format-checked to ensure that the image resolution and color mode meet preset requirements, such as 512×512 resolution and grayscale mode. If the resolution does not meet the requirements, scaling is performed using bilinear interpolation, and finally, the validated image information is output according to the preset format.

[0037] Step S300: The formula text features and formula form features are analyzed to identify the semantic information of the formula structure.

[0038] Specifically, a dual-branch parsing architecture is adopted, which processes the textual features and morphological features of the formula separately. The textual feature branch constructs the grammatical structure of the formula through semantic parsing based on context-free grammar, while the morphological feature branch determines the mathematical formula corresponding to the figure through template matching based on Euclidean distance. Finally, the parsing results of the two branches are merged through semantic fusion to output unified formula structure semantic information.

[0039] For example, in a high school math formula recognition scenario, if the input is an image modality x 2 +y 2 =r 2 The formula text features x, 2 +, y 2 =, r 2 The formula's morphological characteristics are the spatial distribution of characters, such as x. 2 With y 2 The formula is a parallel structure. The text feature branch parses the formula as a quadratic equation in two variables, and the morphological feature branch helps to verify the structural integrity of the formula. The final output is the semantic information of the formula structure of the standard equation of a circle with the origin as the center and r as the radius in a Cartesian coordinate system.

[0040] In one possible implementation, the formula text features and formula form features are parsed separately to identify the formula structural semantic information. Step S300 further includes step S310, which performs semantic parsing on the formula text features to identify the formula structural semantic information. Specifically, a semantic parsing framework based on context-free grammar is adopted, and four steps are executed sequentially: symbol classification, encoding conversion, syntax tree construction, and semantic recognition. First, the symbols in the formula text features are classified according to their types, such as operators, variables, constants, and numbers, and then converted into standard encoding sequences. Then, a formula syntax tree is constructed based on the encoding sequences, and finally, the structural semantic information of the formula is identified based on the syntax tree.

[0041] Step S320 involves performing data formula structure morphology matching and parsing on the formula morphology features, and using the matched mathematical formulas to identify the formula structure semantic information. Specifically, a template library containing various basic mathematical formulas and theorem definitions is constructed. Each template in the template library contains graphic structure features or text spatial distribution features and a corresponding mathematical formula. Then, a similarity calculation algorithm based on Euclidean distance is used to match the formula morphology features with the templates in the template library to determine the most similar basic mathematical formula or theorem definition. Finally, the formula structure semantic information is identified based on the matching results.

[0042] In one possible implementation, semantic parsing is performed on the formula text features to identify the formula structure semantic information. Step S310 further includes step S311, classifying the identified mathematical symbols based on the formula text features. The mathematical symbol types include at least operators, variables, constants, and numbers. Specifically, a symbol classification model combining rule-based and machine learning is adopted. First, a symbol classification rule base is constructed, containing feature descriptions of various types of symbols. For example, the feature of operators is a symbol representing mathematical operations, and the feature of variables is a letter representing an unknown quantity. Then, a support vector machine classifier is used to classify symbols that cannot be classified by the rule base. Finally, the classification results are output.

[0043] Step S312 involves converting various mathematical symbols into standard encoding formats to form a standard encoding sequence. Specifically, a standard encoding dictionary is constructed, containing a unique standard code corresponding to each type of mathematical symbol. The encoding format uses a fixed-length three-digit code. Then, the classified mathematical symbols are sequentially converted into their corresponding standard codes using a lookup table method. Finally, the codes are arranged according to the order of the symbols in the formula to form a standard encoding sequence.

[0044] Step S313: Based on the standard encoding sequence, perform syntax parsing on the mathematical symbols in the formula to construct a formula syntax tree. Specifically, a syntax parsing algorithm based on context-free grammar is adopted. First, the context-free grammar rules of the formula are defined, which include the derivation relationship between non-terminal symbols and terminal symbols. Then, the standard encoding sequence is parsed using recursive descent parsing to construct the formula syntax tree. The nodes of the syntax tree include non-terminal symbol nodes and terminal symbol nodes.

[0045] Step S314: Based on the formula syntax tree, identify the hierarchical structure and operational relationships between mathematical symbol types to generate formula structural semantic information. Specifically, a semantic recognition algorithm based on syntax tree traversal is used. A depth-first traversal algorithm is employed to traverse all nodes of the formula syntax tree. Based on the node type and hierarchical relationship, the type of each mathematical symbol and its semantic role in the formula are identified. For example, the semantic role of a variable is an unknown quantity, and the semantic role of an operator is a mathematical operation. Then, the semantic information of all nodes is integrated to output the structural semantic information of the formula.

[0046] In one possible implementation, the formula morphological features are analyzed using data formula structure morphological matching. The mathematical formulas obtained from the matching analysis are used to identify the semantic information of the formula structure. Step S320 further includes step S321, where, based on the graphical structural features, a search and matching process is performed in a pre-defined mathematical formula template learning space to determine the basic mathematical formulas or theorems that match the current graphical representation. Specifically, a mathematical formula template learning space is constructed, containing a large number of graphical structural feature templates corresponding to basic mathematical formulas and theorems. The features of the templates include the type of graphical representation, the number of vertices, the relationship between edges, angles, etc. A cosine similarity-based retrieval algorithm is used to calculate the similarity between the structural features of the current graphical representation and the templates in the template learning space. Based on the similarity ranking results, the basic mathematical formula or theorem corresponding to the template with the highest similarity is selected.

[0047] Step S322: Based on the annotation mapping relationship, the annotation characters identified in the graphic are substituted as specific parameters into the matching basic mathematical formula or theorem template to instantiate and generate the target mathematical expression corresponding to the current graphic. Specifically, the annotation mapping relationship is parsed to determine the geometric parameters corresponding to the annotation characters in the graphic, such as the side length and angle corresponding to the vertex annotation. Then, parameter placeholders in the basic mathematical formula or theorem template are extracted, such as a, b, and c in the Pythagorean theorem. Through parameter substitution, the geometric parameters corresponding to the annotation characters are substituted into the parameter placeholders to generate the target mathematical expression.

[0048] Step S323 involves performing mathematical symbol recognition and syntactic analysis on the target mathematical expression to clarify the internal operational priorities and structural relationships. The components of the symbolized expression are then mapped back to the original graphical structural features, establishing semantic associations between symbolic elements and geometric primitives to obtain the semantic information of the formula structure. Specifically, rule-based symbol recognition is used to identify the symbols in the target mathematical expression, determine the type of the symbols, set operational priority rules according to mathematical operation rules, clarify the operational priorities and structural relationships of the expression through syntactic analysis (e.g., exponentiation before addition and subtraction), match symbolic elements with geometric primitives in the original graphical structural features through reverse mapping, establish semantic associations, and finally integrate all semantic information to output the semantic information of the formula structure.

[0049] Step S400: Based on the semantic information of the formula structure and combined with a corpus in the mathematical field, perform natural language speech conversion and read the speech information.

[0050] Specifically, the process involves four steps: semantic linearization, corpus optimization, prosodic tagging, and speech synthesis. First, the hierarchical semantic structure is transformed into linear natural language text. Then, a mathematical corpus is used to optimize the text. This corpus contains data such as standard pronunciations of mathematical symbols, common expressions of formulas, and theorem names. Prosodic tags are added to the optimized text, and finally, the prosodic-tagged text is input into a deep learning-based speech synthesis engine to generate speech.

[0051] For example, in the context of speech conversion of university advanced mathematics formulas, the semantic information of the formula structure is: the definite integral of function f on the interval from a to b is equal to F(b) minus F(a). First, it is linearized into: the definite integral of function f on the interval from a to b is equal to F(b) minus F(a) in basic natural language text. Then, a mathematical corpus is used to optimize it into: the definite integral of function f on the closed interval from a to b is equal to the function value of its original function F at point b minus the function value at point a. Then, pause markers are added to the text, with a longer pause after the definite integral and a shorter pause after the equality. Finally, the text with marked rhythm is input into the speech synthesis engine to generate speech information.

[0052] In one possible implementation, natural language speech conversion is performed based on the semantic information of the formula structure combined with a mathematical corpus to read the speech information. Step S400 further includes step S410, which involves traversing the semantic information of the formula structure and linearizing the hierarchical semantic structure into a primary natural language text sequence according to preset mathematical reading rules. The primary natural language text sequence is a structured natural language that conforms to spoken language habits. Specifically, a mathematical reading rule base is constructed, which contains linearization rules for formula semantic structures, such as linearizing an equation structure as: the left expression equals the right expression, and linearizing an exponentiation structure as: the exponent of the base. Then, a depth-first traversal algorithm is used to traverse the hierarchical structure of the formula structure semantic information. According to the rules in the mathematical reading rule base, the semantic information of each level is converted into natural language fragments, and all natural language fragments are spliced ​​together in order to form a primary natural language text sequence.

[0053] Step S420: Invoke the mathematical domain corpus to perform semantic parsing and replacement optimization on the mathematical objects in the primary natural language text sequence, and obtain a natural language description text, including: applying standard reading rules to basic mathematical symbols and function names; matching and applying conventional expression methods or theorem names to formula patterns or geometric relationships; adding descriptive qualifiers to variables and constants according to context information to make the description more complete and unambiguous. Specifically, construct a mathematical domain corpus, which contains data such as standard readings of mathematical symbols, conventional expressions of formulas, theorem names, and descriptive qualifiers for variables and constants. Identify the parts that need to be optimized for the mathematical objects in the primary natural language text sequence through semantic parsing, retrieve the corresponding optimized content in the mathematical domain corpus through retrieval, and finally replace the mathematical objects in the primary natural language text sequence with the optimized content to form an optimized natural language text.

[0054] Step S430: Based on the natural language description text, automatically insert corresponding prosodic marks according to the internal mathematical logical structure, including pauses, stresses, and speech rate adjustments. Specifically, construct a mathematical prosody rule library, which contains rules for inserting prosodic marks according to the mathematical logical structure, such as adding long pauses before and after the equal sign in an equation, adding short pauses before and after the "of" in a power structure, marking stresses on theorem names, and setting the speech rate of complex expressions to be slow. Analyze the mathematical logical structure of the optimized natural language description text through semantic parsing, retrieve the corresponding prosody mark rules in the mathematical prosody rule library through rule matching, and insert the prosody marks into the corresponding positions in the natural language description text.

[0055] Step S440: Input the natural language description text marked with prosody information into a speech synthesis engine to generate speech information. Specifically, use an end-to-end speech synthesis engine, and the model adopted by the engine is a Transformer-based TTS model. This engine includes three core modules: a text encoding layer, an acoustic model layer, and a vocoder layer. Among them, the text encoding layer converts the natural language description text marked with prosody into a vector representation, the acoustic model layer converts the vector representation into Mel spectrum features, and the vocoder layer converts the Mel spectrum features into a speech waveform, and finally outputs speech information.

[0056] In a possible implementation, the method further includes step S500: Match the input formula modality with a preset modality to determine the modality matching relationship. Specifically, construct a preset modality library, which contains three modality types that the system can directly process, namely images, handwritten trajectories, geometric figures, and the feature descriptions of each modality. Extract the feature vector of the input formula modality, calculate the cosine similarity between the features of the input modality and the modality features in the preset modality library, and determine the modality matching relationship according to the similarity result.

[0057] For example, in a scenario where the input formula modality is a speech modality, and there is no speech modality in the preset modality library, the features of the input modality are extracted as one-dimensional waveform data, including audio features. The similarity between this input modality and the image, handwritten trajectory, and geometric shape modalities in the preset modality library are 0.3, 0.2, and 0.1, respectively, indicating a no-match modality relationship. If the input is an image modality, the similarity with the preset image modality is 0.95, indicating a perfect match.

[0058] Step S600: When no modality matching relationship exists, identify the association features and difference features between the current input formula modality and the preset modality. Specifically, construct a modality feature comparison library, which contains association features and difference features between various potential input modalities and preset modalities. For example, the association feature between the speech modality and the image modality is that both contain semantic information of mathematical formulas; the difference feature between the speech modality and the image modality is that the data formats are different—speech data is waveform data, and images are pixel matrices. Extract feature vectors from the current input modality and the preset modality. Determine association features based on the overlap of feature vectors through association degree calculation, and determine difference features based on the difference of feature vectors through difference degree calculation. Output a list of association features and difference features.

[0059] Step S700: Based on the associated features, the data of the input modality is mapped to the feature extraction and recognition process corresponding to the preset modality with the highest correlation. Specifically, a modality correlation calculation model is constructed. This model calculates the correlation between the input modality and each preset modality based on the number and importance weight of associated features, selects the preset modality with the highest correlation and greater than or equal to a threshold, and maps the data of the input modality to the feature extraction and recognition process corresponding to that preset modality.

[0060] Step S800: Based on the aforementioned difference features, the input modality is converted or compensated using an intermediate conversion module to generate an intermediate representation that meets the preset modality processing requirements. Specifically, a corresponding conversion or compensation algorithm is selected according to the type of difference features. For example, for data format differences, a format conversion algorithm, such as speech-to-text or text-to-image, is used; for feature dimension differences, a feature compensation algorithm, such as interpolation completion or feature mapping, is used. The intermediate conversion module executes the selected algorithm to process the input modality and outputs an intermediate representation that meets the preset modality processing requirements. For example, when mapping to an image modality, a 512×512 grayscale image is output.

[0061] Step S900: Using the pre-defined modality recognition process mapped in step S700, mathematical formula recognition processing is performed on the intermediate representation to obtain mathematical semantic information. Specifically, the pre-defined modality recognition process mapped in step S700 is invoked, the intermediate representation generated in step S800 is input into the recognition process, and each core step is executed sequentially according to the recognition logic of the pre-defined modality, finally outputting the processing result, i.e., the mathematical semantic information.

[0062] This application's embodiments acquire arbitrary mathematical formula input modalities from images, handwritten trajectories, and geometric figures. A visual recognition module extracts features from the input information of this modality, obtaining formula text features and formula shape features. These two types of features are then analyzed to identify the structural semantic information of the formula. Combined with a mathematical corpus, the formula structural semantic information is converted into natural language text. Corresponding speech information is then generated through speech synthesis for reading. These techniques solve the technical problems of incomplete multimodal input coverage and low semantic parsing accuracy in existing mathematical formula speech reading technologies, achieving the technical effect of improving the comprehensiveness of multimodal input coverage and the accuracy of formula semantic parsing.

[0063] In the above text, refer to Figure 1 A speech reading method for multimodal mathematical formula recognition according to embodiments of the present invention is described in detail. Next, reference will be made to... Figure 2 A speech reading system for multimodal mathematical formula recognition according to an embodiment of the present invention is described.

[0064] The speech reading system for multimodal mathematical formula recognition according to embodiments of the present invention addresses the technical problems of incomplete multimodal input coverage and low semantic parsing accuracy in existing mathematical formula speech reading technologies, thereby improving the comprehensiveness of multimodal input coverage and the accuracy of formula semantic parsing. The speech reading system for multimodal mathematical formula recognition includes: a mathematical formula input modality acquisition module 10, a visual recognition module 20, a formula structure semantic information recognition module 30, and a natural language speech conversion module 40.

[0065] The mathematical formula input modality acquisition module 10 is used to acquire mathematical formula input modalities, including images, handwritten trajectories, and geometric figures; the visual recognition module 20 is used to extract features from the input multimodal information through the visual recognition module, and recognize the formula text features and formula shape features; the formula structure semantic information recognition module 30 is used to parse the formula text features and formula shape features respectively, and recognize the formula structure semantic information; the natural language speech conversion module 40 is used to perform natural language speech conversion based on the formula structure semantic information and a mathematical domain corpus, and read the speech information.

[0066] The specific configuration of the formula structure semantic information recognition module 30 is described in detail below: As mentioned above, the formula text features and formula form features are analyzed to identify the formula structure semantic information. The formula structure semantic information recognition module 30 may further include: a semantic parsing unit for performing semantic parsing on the formula text features to identify the formula structure semantic information; and a data formula structure form matching parsing unit for performing data formula structure form matching parsing on the formula form features, and using the matched and parsed mathematical formulas to identify the formula structure semantic information.

[0067] The specific configuration of the visual recognition module 20 is described in detail below: As mentioned above, the visual recognition module extracts features from the input multimodal information and identifies formula text features and formula shape features. The visual recognition module 20 may further include: an image conversion unit for performing image conversion on the image, handwritten trajectory, and geometric figure to obtain image information; a standardized symbol recognition unit for performing mathematical character, symbol, and number recognition on the image information through an optical character recognition module to obtain standardized symbols; a label mapping relationship establishment unit for performing edge and graphic structure recognition on the geometric figure and establishing a label mapping relationship between the standardized symbols in the image information and the graphic structure; and a feature acquisition unit for using the standardized symbols as formula text features and the graphic structure features and their label mapping relationships as formula shape features.

[0068] The image conversion unit further includes: a preprocessing subunit for preprocessing the input image, including removing image noise, adjusting image brightness and contrast, rotating and cropping the image to ensure standardization; a smoothing subunit for smoothing the handwritten trajectory input, removing irregular noise, reconstructing the strokes of the trajectory, and converting it into a standardized image format; a structure distribution extraction subunit for edge detection and structure distribution extraction of the geometric shape input, extracting the geometric boundaries of the shape and converting it into standardized image information, wherein geometric annotation mapping relationships are established for shapes with character identifiers; and an image information acquisition subunit for outputting the preprocessed image information, the standardized image information after handwritten trajectory conversion, and the standardized image information after geometric shape conversion to obtain the image information.

[0069] The semantic parsing unit further includes: a mathematical symbol classification subunit for classifying the identified mathematical symbols based on the formula text features, wherein the mathematical symbol types include at least operators, variables, constants, and numbers; a standard encoding format conversion subunit for converting various mathematical symbols into standard encoding formats to form a standard encoding sequence; a syntax parsing subunit for parsing the mathematical symbols in the formula based on the standard encoding sequence to construct a formula syntax tree; and a formula structure semantic information generation subunit for identifying the hierarchical structure and operational relationships between mathematical symbol types based on the formula syntax tree to generate formula structure semantic information.

[0070] The data formula structure morphology matching and parsing unit further includes: a retrieval and matching subunit for searching and matching in a pre-set mathematical formula template learning space based on the graphic structure features to determine the basic mathematical formula or theorem definition that matches the current graphic; a target mathematical expression generation subunit for instantiating and generating a target mathematical expression corresponding to the current graphic by substituting the identified annotation characters in the graphic as specific parameters into the matched basic mathematical formula or theorem template according to the annotation mapping relationship; and a formula structure semantic information acquisition subunit for performing mathematical symbol recognition and syntactic analysis on the target mathematical expression to clarify the internal operation priority and structural relationship, reverse-mapping each component of the symbolic expression back to the original graphic structure features, establishing a semantic association between symbolic elements and geometric primitives, and obtaining the formula structure semantic information.

[0071] The detailed configuration of the natural language speech conversion module 40 is explained below: As mentioned above, based on the semantic information of the formula structure combined with a mathematical corpus, natural language speech conversion is performed to read speech information. The natural language speech conversion module 40 may further include: a primary natural language text sequence generation unit used to traverse the semantic information of the formula structure and, according to preset mathematical reading rules, linearize the hierarchical semantic structure into a primary natural language text sequence, wherein the primary natural language text sequence is a structured natural language that conforms to spoken language habits; and a replacement optimization unit used to call the mathematical corpus and optimize the primary natural language text sequence. The mathematical objects undergo semantic parsing and replacement optimization to obtain natural language description text, including: applying standard reading rules to basic mathematical symbols and function names; matching and applying conventional expressions or theorem names to formula patterns or geometric relationships; adding descriptive qualifiers to variables and constants based on contextual information to make the description more complete and unambiguous; a prosodic marker insertion unit is used to automatically insert corresponding prosodic markers, including pauses, emphasis, and speech rate adjustment, based on the natural language description text and its inherent mathematical logic structure; and a speech synthesis unit is used to input the prosodic-annotated natural language description text into the speech synthesis engine to generate speech information.

[0072] The system may further include: a modality matching module for matching the input formula modality with a preset modality to determine the modality matching relationship; an association difference feature identification module for identifying the association and difference features between the current input formula modality and the preset modality when no modality matching relationship exists; a data mapping module for mapping the data of the input modality to the feature extraction and recognition process corresponding to the preset modality with the highest association degree based on the association features; a data conversion module for performing data conversion or compensation on the input modality based on the difference features through an intermediate conversion module to generate an intermediate representation that meets the processing requirements of the preset modality; and a mathematical formula recognition processing module for performing mathematical formula recognition processing on the intermediate representation using the mapped preset modality recognition process to obtain mathematical semantic information.

[0073] The speech reading system for multimodal mathematical formula recognition provided in this embodiment of the invention can execute the speech reading method for multimodal mathematical formula recognition provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0074] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0075] Based on the foregoing embodiments, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor of an electronic device, can implement the speech reading method for multimodal mathematical formula recognition as described in any of the foregoing embodiments.

[0076] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A speech reading method for multimodal mathematical formula recognition, characterized in that, include: Acquire mathematical formula input modalities, including images, handwritten trajectories, and geometric figures; The visual recognition module extracts features from the input multimodal information to identify formula text features and formula shape features; The formula text features and formula form features are used for analysis to identify the semantic information of the formula structure. Based on the semantic information of the formula structure and combined with a corpus in the field of mathematics, natural language speech conversion is performed to read the speech information.

2. The speech reading method for multimodal mathematical formula recognition according to claim 1, characterized in that, The formula text features and formula form features are analyzed to identify the semantic information of the formula structure, including: Semantic analysis is performed on the textual features of the formula to identify the semantic information of the formula structure; The formula morphological features are analyzed by data formula structure morphological matching, and the mathematical formulas obtained from the matching analysis are used to identify the semantic information of the formula structure.

3. The speech reading method for multimodal mathematical formula recognition according to claim 1, characterized in that, The visual recognition module extracts features from the input multimodal information, recognizing formula text features and formula shape features, including: The image, handwritten trajectory, and geometric figure are recognized and converted to obtain image information; The image information is processed by an optical character recognition module to identify mathematical characters, symbols, and numbers, and standardized symbols are obtained. Edge and graphic structure recognition are performed on the geometric figures, and a mapping relationship between standardized symbols in the image information and graphic structures is established. The standardized symbols are used as formula text features, and the graphic structure features and their annotation mapping relationships are used as formula morphological features.

4. The speech reading method for multimodal mathematical formula recognition according to claim 3, characterized in that, The image, handwritten trajectory, and geometric figure are recognized and converted to obtain image information, including: The input image is preprocessed, including removing image noise, adjusting image brightness and contrast, rotating and cropping the image to ensure image standardization. The handwritten trajectory input is smoothed to remove irregular noise, and the strokes of the trajectory are reconstructed and converted into a standardized image format. Edge detection and structural distribution extraction are performed on the geometric input, the geometric boundaries of the graphic are extracted and converted into standardized image information, wherein geometric annotation mapping relationships are established for those with character identifiers; The image information is obtained by outputting the preprocessed image information, the standardized image information after handwritten trajectory conversion, and the standardized image information after geometric shape conversion.

5. The speech reading method for multimodal mathematical formula recognition according to claim 2, characterized in that, Semantic parsing is performed on the textual features of the formula to identify the semantic information of the formula structure, including: Based on the textual features of the formula, the identified mathematical symbols are classified, and the types of mathematical symbols include at least operators, variables, constants, and numbers; Convert various mathematical symbols into standard encoding formats to form standard encoding sequences; Based on the standard encoding sequence, the mathematical symbols in the formula are parsed to construct a formula syntax tree; Based on the formula syntax tree, the hierarchical structure and operational relationships between mathematical symbol types are identified, and formula structure semantic information is generated.

6. The speech reading method for multimodal mathematical formula recognition according to claim 3, characterized in that, The formula morphological features are analyzed using data formula structure morphological matching. The mathematical formulas analyzed are then used to identify the semantic information of the formula structure, including: Based on the structural features of the graphic, a search and matching process is performed in a pre-set mathematical formula template learning space to determine the basic mathematical formulas or theorems that match the current graphic. Based on the annotation mapping relationship, the annotation characters identified in the graphic are substituted as specific parameters into the matching basic mathematical formula or theorem template to instantiate and generate the target mathematical expression corresponding to the current graphic; The target mathematical expression is subjected to mathematical symbol recognition and syntactic analysis to clarify the internal operation priority and structural relationship. The components of the symbolic expression are mapped back to the original graphic structure features to establish semantic association between symbolic elements and geometric primitives, thereby obtaining the semantic information of the formula structure.

7. The speech reading method for multimodal mathematical formula recognition according to claim 5 or 6, characterized in that, Based on the semantic information of the formula structure combined with a corpus in the mathematical field, natural language speech conversion is performed to read speech information, including: The semantic information of the formula structure is traversed, and the hierarchical semantic structure is linearized into a primary natural language text sequence according to the preset mathematical reading rules. The primary natural language text sequence is a structured natural language that conforms to spoken language habits. The mathematical domain corpus is invoked to perform semantic parsing and replacement optimization on mathematical objects in the primary natural language text sequence, resulting in natural language description text, including: applying standard reading rules to basic mathematical symbols and function names; matching and applying conventional expressions or theorem names to formula patterns or geometric relationships; and adding descriptive qualifiers to variables and constants based on contextual information to make the description more complete and unambiguous. Based on the natural language description text, corresponding prosodic markers, including pauses, emphasis, and speech rate adjustment, are automatically inserted according to the inherent mathematical logic structure. The natural language description text annotated with prosodic information is input into the speech synthesis engine to generate speech information.

8. The speech reading method for multimodal mathematical formula recognition according to claim 7, characterized in that, Also includes: The input formula mode is matched with the preset mode to determine the mode matching relationship; When no modality matching relationship exists, identify the correlation and difference features between the current input formula modality and the preset modality; Based on the aforementioned correlation features, the data of the input modality is mapped to the feature extraction and recognition process corresponding to the preset modality with the highest correlation. Based on the aforementioned differences, the input modality is converted or compensated by an intermediate conversion module to generate an intermediate representation that meets the preset modality processing requirements. Using a pre-defined modality recognition process for mapping, the intermediate representation is processed to identify mathematical formulas and obtain mathematical semantic information.

9. A speech reading system for multimodal mathematical formula recognition, characterized in that, The system is used to implement the speech reading method for multimodal mathematical formula recognition according to any one of claims 1-8, the system comprising: The mathematical formula input modality acquisition module is used to acquire mathematical formula input modalities, including images, handwritten trajectories, and geometric figures. The visual recognition module is used to extract features from the input multimodal information and recognize the text features and shape features of the formula. The formula structure semantic information recognition module is used to analyze the formula text features and formula form features respectively to recognize the formula structure semantic information. The natural language speech conversion module is used to perform natural language speech conversion based on the semantic information of the formula structure combined with a corpus in the field of mathematics, and to read speech information.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the speech reading method for multimodal mathematical formula recognition as described in any one of claims 1-8.