A note keyword recognition method, device and equipment based on a note image

By using an end-to-end note keyword recognition network model, keywords in note images can be quickly identified and parsed, solving the problem that existing tools cannot parse note content in image format, thus improving recognition efficiency and accuracy.

CN117011868BActive Publication Date: 2026-03-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing note-taking tools cannot effectively parse text content in image format, causing users to spend time reading and understanding the notes.

Method used

A note-based keyword recognition method based on note images is adopted. Through an end-to-end network model consisting of a text recognition network, a key semantic extraction network, and a keyword filtering network, a shared contextual semantic recognition layer is used to quickly identify note content and parse key semantics.

Benefits of technology

It improves the efficiency and accuracy of keyword recognition in note images, solves the problem of rapid parsing of note content in image format, and enhances the usability and user experience of note-taking tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011868B_ABST
    Figure CN117011868B_ABST
Patent Text Reader

Abstract

The application discloses a note keyword recognition method based on a note image and a note input device. The embodiment of the application can be applied to an artificial intelligence scene. Specifically, the method comprises the following steps: collecting a note image of target note information; inputting the note image into a text recognition network in a note keyword recognition network to perform text recognition, and obtaining note recognition text; inputting the note recognition text into a key semantic extraction network in the note keyword recognition network to perform key semantic extraction, and obtaining target semantic feature information; inputting the target semantic feature information and multiple candidate keywords of the note recognition text into a keyword screening network in the note keyword recognition network to perform keyword screening, and determining a note keyword in the multiple candidate keywords; wherein, the text recognition network and the keyword screening network share a context semantic recognition layer. By using the technical scheme provided in the application, the note image can be recognized in real time, and the keyword recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of software testing, in particular to a note keyword recognition method and device based on note images, a note keyword recognition equipment and a note recording equipment. BACKGROUND

[0002] Note recording is an indispensable content storage method in our daily life and work. With the rapid development of the economy and society and the evolution of education methods, a note recording method using pictures as carriers has emerged.

[0003] However, most of the current note recording tools can only directly save the text in picture format or save the original format files of various web pages through related plug-ins, but cannot perform key semantic analysis on the content in the text in picture format, so that the text reader needs to spend time and effort to read and understand the note content in the picture. SUMMARY

[0004] The present application provides a note keyword recognition method and device based on note images, equipment and storage medium, which can reduce the data size of the network model while quickly identifying the note content and quickly analyzing the key semantics of the note pictures, thereby improving the efficiency and accuracy of note keyword recognition of note pictures. The technical solutions of the present application are as follows:

[0005] On the one hand, a note keyword recognition method based on note images is provided, which comprises:

[0006] Collecting a note image of target note information;

[0007] Inputting the note image into a text recognition network in a note keyword recognition network for text recognition to obtain note recognition text corresponding to the note image;

[0008] Inputting the note recognition text into a key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain target semantic feature information;

[0009] Obtaining a plurality of candidate keywords corresponding to the note recognition text;

[0010] Inputting the target semantic feature information and the plurality of candidate keywords into a keyword screening network in the note keyword recognition network for keyword screening to determine a note keyword in the plurality of candidate keywords;

[0011] The text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is used for context semantic recognition of feature information.

[0012] In another aspect, a pen recording device is provided, the pen recording device comprising: an image acquisition module and a processor, wherein:

[0013] The image acquisition module is configured to acquire a note image of target note information.

[0014] The processor is configured to perform text recognition on the note image by using a text recognition network in a note keyword recognition network to obtain note recognition text corresponding to the note image, perform key semantic extraction on the note recognition text by using a key semantic extraction network in the note keyword recognition network to obtain target semantic feature information, acquire a plurality of candidate keywords corresponding to the note recognition text, and perform keyword screening on the target semantic feature information and the plurality of candidate keywords by using a keyword screening network in the note keyword recognition network to determine a note keyword in the plurality of candidate keywords. The text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on feature information.

[0015] In another aspect, a pen note keyword recognition device based on a note image is provided, the device comprising:

[0016] An image acquisition module is configured to acquire a note image of target note information.

[0017] A text recognition module is configured to perform text recognition on the note image by using a text recognition network in a note keyword recognition network to obtain note recognition text corresponding to the note image.

[0018] A key semantic extraction module is configured to perform key semantic extraction on the note recognition text by using a key semantic extraction network in the note keyword recognition network to obtain target semantic feature information.

[0019] A candidate keyword acquisition module is configured to acquire a plurality of candidate keywords corresponding to the note recognition text.

[0020] A keyword screening module is configured to perform keyword screening on the target semantic feature information and the plurality of candidate keywords by using a keyword screening network in the note keyword recognition network to determine a note keyword in the plurality of candidate keywords.

[0021] The text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on feature information.

[0022] In another aspect, a note keyword recognition device based on a note image is provided, the device comprising a processor and a memory, the memory having stored therein at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the note keyword recognition method based on a note image as described in the first aspect.

[0023] In another aspect, a computer-readable storage medium is provided, the storage medium having stored therein at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the note keyword recognition method based on a note image as described in the first aspect.

[0024] In another aspect, a computer program product or computer program is provided, the computer program product or computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the note keyword recognition method based on a note image as described in the first aspect.

[0025] The note keyword recognition method, device, equipment and storage medium based on a note image provided by the present application have the following technical effects:

[0026] In the application scenario of recognizing and processing a text image, based on a text recognition network and a keyword screening network of a shared context semantic recognition layer, an end-to-end note keyword recognition network is generated, which can reduce the data size of the network model. By inputting a note image containing target note information collected into the text recognition network in the note keyword recognition network for text recognition, note recognition text corresponding to the note image is obtained. Then, the note recognition text is input into the key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain target semantic feature information. Then, the target semantic feature information and a plurality of candidate keywords corresponding to the note recognition text are input into the keyword screening network in the note keyword recognition network for keyword screening to determine note keywords in the plurality of candidate keywords. Not only can the note content of the note in a picture format be quickly recognized, but also the key semantics of the note content can be quickly analyzed and extracted, which improves the efficiency and accuracy of note keyword recognition of the note image, solves the problem that the current note recognition tool cannot fully analyze the key content of the note in a picture format, and greatly improves the use effect and user experience of the note recognition tool. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, the drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these drawings without creative labor.

[0028] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0029] Figure 2 is a flowchart of a note keyword recognition method based on a note image provided by an embodiment of the present application;

[0030] Figure 3 is a flowchart of a note image for collecting target note information provided by an embodiment of the present application;

[0031] Figure 4 is a network structure diagram of a note keyword recognition network provided by an embodiment of the present application;

[0032] Figure 5 is a flowchart of inputting a note image into a text recognition network in a note keyword recognition network for text recognition, to obtain note recognition text corresponding to the note image provided by an embodiment of the present application;

[0033] Figure 6 is a flowchart of inputting note recognition text into a key semantic extraction network in a note keyword recognition network for key semantic extraction, to obtain target semantic feature information provided by an embodiment of the present application;

[0034] Figure 7 is a flowchart of inputting topic keyword feature information and high-frequency semantic feature information into a semantic fusion network for semantic fusion, to obtain target semantic feature information provided by an embodiment of the present application;

[0035] Figure 8 is a flowchart of inputting target semantic feature information and a plurality of candidate keywords into a keyword screening network in a note keyword recognition network for keyword screening, to determine note keywords in the plurality of candidate keywords provided by an embodiment of the present application;

[0036] Figure 9 is a flowchart of a note keyword recognition network training method provided by an embodiment of the present application;

[0037] Figure 10 is another network structure diagram provided by an embodiment of the present application;

[0038] Figure 11is a component block diagram of a note keyword recognition device based on a note image provided by an embodiment of the present application.

[0039] Figure 12 is a structural schematic diagram of a note keyword recognition device based on a note image provided by an embodiment of the present application. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0041] It should be noted that the terms "include" and "have" and any variations thereof in the specification and claims of the present application and the above-described drawings are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units need not be limited to those clearly listed steps or units, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0042] It can be understood that in the specific embodiments of the present application, data related to user information is involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0043] Please refer to Figure 1 , Figure 1is a schematic diagram of an application environment provided by an embodiment of the present application. The application environment can include a client 10 and a server 20. The client 10 and the server 20 can be indirectly connected through a wireless communication mode. A user can initiate a note keyword recognition instruction for a note image including target note content at the client 10. The client 10 sends the note keyword recognition instruction to the server 20. The server 20 responds to the note keyword recognition instruction, first inputs the note image into a text recognition network in a note keyword recognition network for text recognition to obtain note recognition text corresponding to the note image, then inputs the note recognition text into a key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain target semantic feature information, then acquires a plurality of candidate keywords corresponding to the note recognition text, then inputs the target semantic feature information and the plurality of candidate keywords into a keyword screening network in the note keyword recognition network for keyword screening to determine a note keyword in the plurality of candidate keywords, and returns the note recognition text and the note keyword to the client 10, so that the client 10 displays the note recognition text and the note keyword. The text recognition network and the keyword screening network share a context semantic recognition layer. The context semantic recognition layer is used for context semantic recognition of feature information. It should be noted that, Figure 1 is merely an example.

[0044] In addition, as an optional implementation, the note keyword recognition method based on a note image provided by the present application can also be applied in, but is not limited to, an application environment as shown in Figure 1 , such as in a client. A user can initiate a note keyword recognition instruction for a note image including target note content at the client. The client responds to the recognition processing instruction, first inputs the note image into a text recognition network in a note keyword recognition network for text recognition to obtain note recognition text corresponding to the note image, then inputs the note recognition text into a key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain target semantic feature information, then acquires a plurality of candidate keywords corresponding to the note recognition text, then inputs the target semantic feature information and the plurality of candidate keywords into a keyword screening network in the note keyword recognition network for keyword screening to determine a note keyword in the plurality of candidate keywords, and displays the note recognition text and the note keyword.

[0045] The client can be an entity device such as a smart phone, a computer (e.g., a desktop computer, a tablet computer, a notebook computer), a digital assistant, a smart voice interaction device (e.g., a smart speaker), a smart wearable device, a vehicle terminal, etc., or a software such as a computer program running in the entity device. The operating system corresponding to the first client can be an Android system, an iOS system (a mobile operating system developed by Apple Inc.), a Linux system, a Microsoft Windows system, etc.

[0046] The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The server can include a network communication unit, a processor, a memory, etc. The server can provide background services for the corresponding client.

[0047] The client 10 and the server 20 described above can be used to build a system for identifying note keywords from a note image. The system can be a distributed system. For example, the distributed system is a blockchain system formed by multiple nodes (any form of computing device connected to a network, such as a server or a user terminal) and clients. The nodes form a P2P (Peer To Peer) network, and the P2P protocol is an application layer protocol running on the TCP (Transmission Control Protocol) protocol. In the distributed system, any machine such as a server or a terminal can join as a node, and the node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer.

[0048] The functions of the nodes in the blockchain system described above include:

[0049] 1) Routing, a basic function of the node, used to support communication between nodes.

[0050] In addition to the routing function, the node can have the following functions:

[0051] 2) Application, for deployment in a blockchain, to implement a specific business according to actual business needs, record data related to the implementation function to form record data, carry a digital signature in the record data to represent the source of the task data, send the record data to other nodes in the blockchain system, and add the record data to the temporary block when the other nodes verify the source and integrity of the record data successfully.

[0052] 3) Blockchain, including a series of blocks connected to each other in chronological order, once a new block is added to the blockchain, it will not be removed, and the block records the record data submitted by the nodes in the blockchain system.

[0053] It should be noted that the note keyword recognition method based on note images provided in the present application can be applied to the client side or the server side, and is not limited to the above application environment embodiments.

[0054] The following describes a specific embodiment of a note keyword recognition method based on note images provided by the present application, Figure 2 is a flowchart of a note keyword recognition method based on note images provided by an embodiment of the present application, and the present application provides a method operation step as described in the embodiment or flowchart, but more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiment is only one of the many execution orders, and does not represent the only execution order. In actual system or product execution, the method order shown in the embodiment or the drawing can be executed in sequence or in parallel (for example, parallel processor or multi-threaded processing environment). Specifically, as Figure 2 shown, the method can include:

[0055] S201, collecting a note image of target note information.

[0056] In an embodiment of the present specification, the target note information can represent an information carrier containing note content, and the note content can be information annotated to facilitate understanding of the content of a certain resource. Specifically, the target note information can include but is not limited to: PPT display pages, videos, book pages, etc. containing note content.

[0057] In an embodiment of the present specification, the note image can be an image obtained by photographing the target note information, for example, the note image can include: PPT photograph, book photograph, video frame image, etc.

[0058] In one specific embodiment, the note image can be an image obtained by real-time photographing of the target note information by a pen recording device. Optionally, the pen recording device can include: a PPT recording device and a book recording device, etc.

[0059] In practical applications, there are usually many interferences and noises on the photographed or stored image content containing note information, which cannot be directly subjected to text detection and recognition, and thus the note image needs to be preprocessed, and the preprocessing steps include graying, binarization, tilt detection and correction, image smoothing, normalization, etc.

[0060] In the embodiments of the present specification, as shown in Figure 3 The note image collecting the target note information can include:

[0061] S301, collecting a to-be-processed note image of target note information.

[0062] Specifically, the to-be-processed note image can be an unprocessed note image directly photographed or stored.

[0063] S302, performing graying processing on the to-be-processed note image to obtain a first note image.

[0064] In practical applications, graying processing is to map the originally three-dimensional described pixel points in the image to one-dimensional described pixel points, so as to filter out the interference information of text recognition.

[0065] Specifically, there are the following several ways for image graying processing:

[0066] 1) The brightness of RGB three components in the color image is taken as the gray value of the three gray images, and one gray image can be selected according to the application needs.

[0067] Gray1(i, j) = R(i, j)

[0068] Gray2(i, j) = G(i, j)

[0069] Gray3(i, j) = B(i, j)

[0070] Wherein, (i, j) represents the pixel point coordinates in the image, R(i, j) represents the R component brightness of the pixel point with coordinates (i, j), G(i, j) represents the G component brightness of the pixel point with coordinates (i, j), B(i, j) represents the B component brightness of the pixel point with coordinates (i, j), and Gray1(i, j), Gray2(i, j), and Gray3(i, j) represent the gray values of the three gray images, respectively.

[0071] 2) The maximum value of the brightness of RGB three components in the color image is taken as the gray value of the gray image.

[0072] Gray(i, j) = max{R(i, j), G(i, j) , B(i, j)}

[0073] Wherein, (i, j) represents the pixel point coordinates in the image, R(i, j) represents the R component brightness of the pixel point with coordinates (i, j), G(i, j) represents the G component brightness of the pixel point with coordinates (i, j), B(i, j) represents the B component brightness of the pixel point with coordinates (i, j), and Gray(i, j) represents the gray value of the gray scale image.

[0074] 3) The average value of the RGB three component brightness in the color image is taken as the gray value of the gray scale image.

[0075] Gray(i, j) = (R(i, j) + G(i, j) + B(i, j)) / 3

[0076] Wherein, (i, j) represents the pixel point coordinates in the image, R(i, j) represents the R component brightness of the pixel point with coordinates (i, j), G(i, j) represents the G component brightness of the pixel point with coordinates (i, j), B(i, j) represents the B component brightness of the pixel point with coordinates (i, j), and Gray(i, j) represents the gray value of the gray scale image.

[0077] S303, the first note image is binarized to obtain a second note image.

[0078] Specifically, the binarization processing is to convert the gray value image signal into a binary image signal with only black (l) and white (0), so as to further separate the text from the background. The most commonly used method for gray scale image binarization is threshold method.

[0079] S304, the second note image is subjected to inclination correction processing to obtain a third note image.

[0080] Specifically, the inclination correction processing on the second note image can include:

[0081] 1) Based on the projection graph method, the image is projected along different directions. When the projection direction and the text line direction are consistent, the peak value of the text line on the projection graph is the largest, and there is an obvious peak and valley in the projection graph. At this time, the projection direction is the inclination angle, and the image is corrected based on the inclination angle.

[0082] 2) The method based on Hough transform is to map the foreground pixels in the image to the polar coordinate space by using the characteristics of Hough transform. The inclination angle of the document image is obtained by counting the cumulative values of the points in the polar coordinate space. The image is corrected based on the inclination angle. Specifically, the characteristics of Hough transform is to transform the curve (including straight line) in the image space to the parameter space. By detecting the extreme points in the parameter space, the description parameters of the curve are determined, so as to extract the regular curve in the image.

[0083] 3) Based on the nearest neighbor clustering method, the center point of the character connected domain in any sub-region of the image is taken as a feature point, the continuity of the points on the baseline is used to calculate the direction angle of the corresponding text line, so as to obtain the tilt angle of the entire text page in the image, and the image is corrected based on the tilt angle.

[0084] S305, the third note image is subjected to noise filtering processing to obtain a fourth note image.

[0085] Specifically, the noise filtering processing can include mean filtering, box filtering, Gaussian filtering, median filtering, bilateral filtering, etc.

[0086] S306, the fourth note image is subjected to normalization processing to obtain a note image.

[0087] Specifically, the normalization processing is to process the text of any size in the input image into standard text of a uniform size. The normalization processing can include position normalization and font size normalization, the position normalization can include centroid-based position normalization and text outer frame-based position normalization; the font size normalization can linearly scale the text outer frame to a standard size text in proportion, or can normalize the size according to the distribution of text black pixels in the horizontal and vertical directions.

[0088] As can be seen from the above embodiments, by performing various pre-processing on the image containing note information, the interference information in the image can be filtered out, thereby improving the accuracy of subsequent note text recognition.

[0089] S202, the note image is input into a text recognition network in the note keyword recognition network for text recognition to obtain note recognition text corresponding to the note image.

[0090] In the embodiments of the present specification, the note keyword recognition network can be used for note keyword recognition on the note content in the note image.

[0091] In one specific embodiment, the note keyword recognition network can be obtained by pre-training a preset note keyword recognition network based on sample note images with note keyword annotations.

[0092] In one specific embodiment, as shown in Figure 4 The note keyword recognition network can be an end-to-end network model including a text recognition network, a key semantic extraction network, and a keyword screening network, wherein the text recognition network and the keyword screening network can share a context semantic recognition layer, and the context semantic recognition layer is used for context semantic recognition on feature information.

[0093] In one specific embodiment, the text recognition network can further include an image feature extraction layer and a text transcription layer, as shown in Figure 5 The text recognition by inputting the note image into the text recognition network in the note keyword recognition network to obtain the note recognition text corresponding to the note image can include:

[0094] S501, inputting the note image into the image feature extraction layer to perform image feature extraction to obtain an image feature sequence.

[0095] In one specific embodiment, the image feature sequence can represent a plurality of ordered feature maps of the note image.

[0096] In one specific embodiment, the image feature extraction layer can be a convolutional layer, and correspondingly, the image feature extraction by inputting the note image into the image feature extraction layer to obtain the image feature sequence can include: performing feature extraction on the note image based on the convolutional layer to obtain a plurality of convolutional features corresponding to the note image, and through sequential processing on the plurality of convolutional features corresponding to the note image, an image feature sequence corresponding to the note image can be obtained.

[0097] Optionally, the convolutional layer can include, but is not limited to, a deep learning network layer such as a convolutional neural network, a You Only Look Once (YOLO), a Single Shot MultiBox Detector (SSD), etc.

[0098] S502, inputting the image feature sequence into the context semantic recognition layer to perform context semantic recognition to obtain character feature information corresponding to the image feature sequence.

[0099] In one specific embodiment, the character feature information can represent a character distribution probability corresponding to each image feature in the image feature sequence.

[0100] In one specific embodiment, the context semantic recognition layer can be a recurrent layer, and correspondingly, the context semantic recognition by inputting the image feature sequence into the context semantic recognition layer to obtain the character feature information corresponding to the image feature sequence can include: performing recognition processing on the image feature sequence based on the recurrent layer to determine the character feature information corresponding to the image feature sequence.

[0101] Optionally, the recurrent layer can include, but is not limited to, a BI-LSTM (Bidirectional Long Short-Term Memory) layer, a recurrent neural network layer, or other deep learning network layers.

[0102] S503, inputting the character feature information into the text transcription layer to perform feature conversion to obtain the note recognition text.

[0103] In a specific embodiment, the inputting the character feature information into the text transcription layer for feature conversion to obtain the note recognition text can include: performing feature conversion on the character feature information based on the text transcription layer, and integrating empty characters and repeated characters in the character features obtained after the feature conversion to obtain the note recognition text of the note image.

[0104] Optionally, the text transcription layer can include a network layer constructed by, but not limited to, a Connectionist Temporal Classification (CTC) algorithm or other algorithms.

[0105] As can be seen from the above embodiments, the note text recognition is performed on the note image by the text recognition network including the image feature extraction layer, the context semantic recognition layer, and the text transcription layer, which can quickly recognize the note text content while improving the accuracy of the note text recognition.

[0106] S203, inputting the note recognition text into a key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain target semantic feature information.

[0107] In the embodiments of the present disclosure, the target semantic feature information can represent aggregated semantic information of a plurality of text keywords in the note recognition text. In a specific embodiment, the target semantic feature information can be in the form of a target semantic feature vector.

[0108] In a specific embodiment, the key semantic extraction network can include a topic extraction network, a high-frequency semantic mining network, and a semantic fusion network, as shown in Figure 6 The inputting the note recognition text into the key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain the target semantic feature information can include:

[0109] S601, inputting the note recognition text into a topic extraction network for topic extraction to obtain topic keyword feature information.

[0110] In a specific embodiment, the topic keyword feature information can represent feature information of keywords related to the topic of the full text of the note recognition text. The topic keyword feature information can be in the form of a topic keyword feature vector.

[0111] In a specific embodiment, the topic extraction network can be used for topic extraction on the note recognition text. Optionally, the topic extraction network can include, but not limited to, an LDA (Latent Dirichlet Allocation) topic network, a PLSA (Probabilistic Latent Semantic Analysis) topic network, etc.

[0112] Taking the LDA topic network as an example, the LDA topic network is a statistical network model used to find a set of potential topics containing a specific probability, such as “economy”, “transportation”, etc., from a document set, forming a three-layer structure of word-topic-text content. The characteristics of the topic are represented by the distribution of words, reflecting the topic distribution of the text content.

[0113] The LDA topic network can be represented as:

[0114] wherein, tp(w i |d k ) represents a text-topic distribution matrix, p() represents a probability, d k represents text content, w i represents a word in the text content d k , and t j represents a topic implied in the text content.

[0115] Specifically, the LDA topic network selects a certain topic corresponding to the text content with a certain probability, and then selects a certain word from the topic with a certain probability. These two steps are repeated until the entire text is generated. In the case of network convergence, the topic distribution characteristic information based on the text content can be obtained. The topic distribution characteristic information can represent the probability of the text where the word belongs to each topic word. Words with similar semantics have similar topic distribution characteristic information, and thus the topic distribution characteristic information of the current note recognition text is obtained.

[0116] In one specific embodiment, the above-mentioned inputting the note recognition text into the topic extraction network for topic extraction to obtain the topic word characteristic information can include: inputting the note recognition text into the topic extraction network for topic extraction to obtain the topic distribution characteristic information of the note recognition text, then performing word segmentation processing on the note recognition text to obtain a plurality of segmented words, obtaining the word characteristic information of each segmented word based on the trained word feature extraction network, performing similarity analysis on the word characteristic information of each segmented word and the topic distribution characteristic information of the note recognition text to obtain the characteristic similarity corresponding to the word characteristic information of each segmented word, obtaining the first preset number of similar word characteristic information of the topic distribution characteristic information according to the characteristic similarity from large to small, and taking the first preset number of similar word characteristic information as the topic word characteristic information.

[0117] In a specific embodiment, the word feature information can include a word vector, the topic distribution feature information can include a topic distribution vector, and the feature similarity can include first vector distance information. Accordingly, the similarity analysis of the word feature information of each segmented word and the topic distribution feature information of the note recognition text can include calculating the first vector distance information of the word vector of each segmented word and the topic distribution vector of the note recognition text.

[0118] In a specific embodiment, the first vector distance information can include a cosine distance, an Euclidean distance, a Manhattan distance, or the like.

[0119] In actual applications, the first preset number can be set in advance according to the accuracy of topic word extraction.

[0120] S602, inputting the note recognition text into a high-frequency semantic mining network to perform high-frequency semantic mining, and obtaining high-frequency semantic feature information.

[0121] In a specific embodiment, the high-frequency semantic feature information can represent the feature information of a keyword with a frequency greater than a preset frequency condition in the note recognition text. Specifically, the preset frequency condition can be set in advance in combination with the number of sentences of the note recognition text and the accuracy of keyword recognition in actual applications.

[0122] In a specific embodiment, the high-frequency semantic mining network can be used for high-frequency semantic mining of the note recognition text. Optionally, the high-frequency semantic mining network can include, but is not limited to, a high-frequency semantic mining network based on a PrefixSpan algorithm (Prefix-projected Sequential pattern mining, prefix-projected sequential pattern mining algorithm), a high-frequency semantic mining network based on a FreeSpan algorithm (Frequent pattern-projected Sequential pattern mining, frequent pattern-projected sequential pattern mining algorithm), or the like.

[0123] Taking the high-frequency semantic mining network based on the PrefixSpan algorithm as an example, the PrefixSpan algorithm can mine frequent word sequences of various lengths that meet a minimum support threshold in the local context of the note recognition text, take the frequent word sequences as high-frequency keywords, obtain high-frequency keyword feature information corresponding to the high-frequency keywords based on the trained word feature extraction network, and obtain the high-frequency semantic feature information based on the high-frequency keyword feature information.

[0124] Specifically, the minimum support threshold can represent a minimum number threshold of the sentences in which the corresponding frequent word sequence appears in the segmented sentences of the note recognition text. In a specific embodiment, the minimum support threshold = a x n, where n represents the number of segmented sentences obtained after the note recognition text is segmented, and a represents a preset semantic support degree. Specifically, the preset semantic support degree can represent a minimum proportion of the number of sentences in which the corresponding frequent word sequence appears in the segmented sentences of the note recognition text. The preset semantic support degree can be adjusted according to the number of segmented sentences of the note recognition text in actual application.

[0125] In an optional embodiment, the semantic support degree of the high-frequency keyword feature information can be obtained based on the number of sentences in which the high-frequency keyword corresponding to the high-frequency keyword feature information appears in the note recognition text. Specifically, the semantic support degree can represent the frequency of the high-frequency keyword corresponding to the high-frequency keyword feature information appearing in the note recognition text.

[0126] In S603, the topic word feature information and the high-frequency semantic feature information are input into a semantic fusion network for semantic fusion to obtain target semantic feature information.

[0127] In a specific embodiment, the high-frequency semantic feature information can include at least one high-frequency keyword feature information and a semantic support degree corresponding to each high-frequency keyword feature information. The semantic fusion network includes a weighting layer and a fusion layer, as shown in Figure 7 The above semantic fusion of the topic word feature information and the high-frequency semantic feature information in the semantic fusion network to obtain the target semantic feature information can include:

[0128] In S701, the semantic support degree and the at least one high-frequency keyword feature information are input into the weighting layer. The at least one high-frequency keyword feature information is weighted based on the semantic support degree to obtain initial semantic feature information.

[0129] In S702, the initial semantic feature information and the topic word feature information are input into the fusion layer for semantic fusion to obtain target semantic feature information.

[0130] In a specific embodiment, the high-frequency keyword feature information, the initial semantic feature information, the topic word feature information, and the target semantic feature information can be in the form of feature vectors. Specifically, the high-frequency keyword feature vector can be obtained by averaging the word vectors of multiple words in the high-frequency keyword, and the topic word feature vector can be obtained by averaging the word vectors of multiple topic words in the note recognition text.

[0131] Taking the high-frequency keyword feature vector [a, b, c] and the semantic support degree corresponding to the high-frequency keyword feature vector being 0.85 as an example, an initial semantic feature vector is [0.85a, 0.85b, 0.85c], and assuming that the topic keyword feature information is [A, B, C], a target semantic feature vector is [0.85a + A, 0.85b + B, 0.85c + C].

[0132] As can be seen from the above embodiments, not only the global topic information of the note content is taken into account, but also the locally frequently occurring feature information is mined through sequence pattern mining, and the global topic information and the locally frequently occurring sequence pattern information are fused, so that more dimensional word information reference for keyword extraction can be provided, and the accuracy of keyword semantic extraction is improved.

[0133] In S204, a plurality of candidate keywords corresponding to the note recognition text are obtained.

[0134] In an optional embodiment, the above obtaining the plurality of candidate keywords corresponding to the note recognition text can include: performing a word segmentation processing on the note recognition text to obtain a plurality of segmented words, and taking the plurality of segmented words as the plurality of candidate keywords.

[0135] In another optional embodiment, the above obtaining the plurality of candidate keywords corresponding to the note recognition text can include: taking the topic keywords corresponding to the topic keyword feature information and the high-frequency keywords corresponding to the high-frequency semantic feature information as the plurality of candidate keywords.

[0136] In S205, the target semantic feature information and the plurality of candidate keywords are input into a keyword screening network in the note keyword recognition network for keyword screening to determine a note keyword in the plurality of candidate keywords; wherein the text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on the feature information.

[0137] In a specific embodiment, the note keyword can be a preferred text keyword corresponding to the note recognition text of the note image.

[0138] In a specific embodiment, the keyword screening network further includes an association analysis layer and a keyword screening layer, as shown in FIG. 8. Figure 8 The above inputting the target semantic feature information and the plurality of candidate keywords into the keyword screening network in the note keyword recognition network for keyword screening to determine the note keyword in the plurality of candidate keywords can include:

[0139] In S801, the target semantic feature information is input into the context semantic recognition layer for context semantic recognition to obtain a context semantic feature.

[0140] In a specific embodiment, the context semantic feature can represent feature information obtained after context semantic encoding of a plurality of text keywords in the recognized text of the note.

[0141] In a specific embodiment, the context semantic recognition layer can be a recurrent layer, and the context semantic feature can be obtained by inputting the target semantic feature information into the context semantic recognition layer for context semantic recognition, including: performing context encoding processing on the target semantic feature information based on the recurrent layer to obtain the context semantic feature.

[0142] Optionally, the recurrent layer can include, but is not limited to, a bidirectional long short-term memory network layer (BI-LSTM), a recurrent neural network layer, or other deep learning network layers.

[0143] S802, input the context semantic feature and the plurality of candidate keywords into the association analysis layer for feature association analysis to obtain feature association information between each candidate keyword and the context semantic feature.

[0144] In a specific embodiment, the feature association information can represent a similarity between each candidate keyword and the context semantic feature.

[0145] In a specific embodiment, the context semantic feature can include a context semantic feature vector, and the feature association information can include second vector distance information, and the above-mentioned inputting the context semantic feature and the plurality of candidate keywords into the association analysis layer for feature association analysis to obtain the feature association information between each candidate keyword and the context semantic feature includes: performing word feature extraction on the plurality of candidate keywords to obtain a plurality of candidate keyword feature vectors corresponding to the plurality of candidate keywords respectively, and calculating second vector distance information between each candidate keyword feature vector and the context semantic feature vector.

[0146] In a specific embodiment, the second vector distance information can include an angle distance, a cosine distance, an Euclidean distance, and a Manhattan distance, etc.

[0147] In a specific embodiment, the association analysis layer can include, but is not limited to, an AM-Softmax (Additive Margin Softmax, an angle loss function) layer, a Sigmoid function layer.

[0148] S803, input the feature association information into the keyword screening layer for keyword screening to determine the note keyword.

[0149] Specifically, based on the similarity represented by the feature association information, a second preset number of similar candidate keywords of the context semantic feature are obtained from large to small, and the second preset number of similar candidate keywords are taken as the note keyword.

[0150] Taking the keyword screening network including a BI-LSTM layer, an AM-Softmax layer and a ranking output layer as an example, the input target semantic feature information of the keyword screening network can include a target semantic feature vector. The BI-LSTM layer can perform context encoding processing on the target semantic feature vector to obtain a context semantic feature vector, that is, y=Bi-LSTM(x), x is the target semantic feature vector, and y is the context semantic feature vector output by the BI-LSTM layer. The AM-Softmax layer can calculate the angle distance between each candidate word feature vector in the candidate word feature vector corresponding to the multiple candidate keywords of the note recognition text and the context semantic feature vector, and take the angle distance as the probability that the candidate keyword corresponding to the candidate word feature vector belongs to the note keyword, that is, p=am-softmax(<y,c1>,<y,c2>,...,<y,c n ), that is, p=am-softmax(<y,c1>,<y,c2>,...,<y,c n >). The ranking output layer can determine the second preset number of candidate keywords in the multiple candidate keywords according to the order from large to small of the probability.

[0151] As can be seen from the above embodiments, by inputting the target semantic feature information of the note recognition text and the multiple candidate keywords into the keyword screening network including the context semantic recognition layer, the association analysis layer and the keyword screening layer for keyword screening, the accuracy of note keyword screening can be improved.

[0152] In the embodiments of the present specification, as shown in Figure 9 , the above note keyword recognition network is trained in the following manner:

[0153] S901, obtaining a sample note image and a labeled note keyword corresponding to the sample note image.

[0154] In actual application, before network training, training data can be determined first. Specifically, in the embodiments of the present application, a sample note image containing a labeled note keyword can be obtained as training data.

[0155] Specifically, the labeled note keyword can be a preset note keyword label pre-labeled for the sample note image.

[0156] S902, inputting the sample note image into a preset text recognition network in the preset note keyword recognition network for text recognition to obtain a sample note recognition text corresponding to the sample note image.

[0157] S903, input the sample note recognition text into a preset key semantic extraction network in the preset note keyword recognition network for key semantic extraction, to obtain sample semantic feature information.

[0158] S904, obtain a plurality of sample candidate keywords corresponding to the sample note recognition text.

[0159] S905, input the sample semantic feature information and the plurality of sample candidate keywords into a preset keyword screening network in the preset note keyword recognition network for keyword screening, to determine a sample note keyword in the plurality of sample candidate keywords.

[0160] S906, based on the labeled note keyword and the sample note keyword, determine target loss information.

[0161] S907, based on the target loss information, train the preset note keyword recognition network to generate a note keyword recognition network.

[0162] Among them, the preset text recognition network and the preset keyword screening network share a preset context semantic recognition layer, and the preset context semantic recognition layer is used for context semantic recognition of sample feature information.

[0163] In an optional embodiment, the above-mentioned sample note keyword can include a sample note keyword label of a sample note image, and correspondingly, the above-mentioned target loss information can include a keyword label loss;

[0164] Correspondingly, the above-mentioned determination of target loss information based on the labeled note keyword and the sample note keyword can include:

[0165] According to the preset note keyword label and the sample note keyword label, determine the keyword label loss.

[0166] In a specific embodiment, the above-mentioned determination of keyword label loss according to the preset note keyword label and the sample note keyword label can include determination of keyword label loss between the preset note keyword label and the sample note keyword label based on a preset loss function.

[0167] In a specific embodiment, the keyword label loss can represent the difference between the preset note keyword label and the sample note keyword label.

[0168] In a specific embodiment, the preset loss function can include but is not limited to Margin loss function, cross-entropy loss function, logical loss function, exponential loss function, etc.

[0169] In an optional embodiment, based on the target loss information, training the preset note keyword recognition network to generate a note keyword recognition network can include:

[0170] S9071, update the network parameters of the preset note keyword recognition network based on the target loss information.

[0171] Specifically, updating the network parameters of the preset note keyword recognition network can include updating the network parameters of the preset text recognition network, updating the network parameters of the preset key semantic extraction network, and updating the network parameters of the preset keyword screening network.

[0172] S9072, based on the updated preset note keyword recognition network, repeatedly perform the note keyword recognition training iteration operation including steps S902-S906 and S9071 until the note keyword recognition convergence condition is reached.

[0173] S9073, take the preset note keyword recognition network obtained when the note keyword recognition convergence condition is reached as the note keyword recognition network.

[0174] In an optional embodiment, the note keyword recognition convergence condition reached above can be that the number of training iteration operations reaches a preset training number. Optionally, the note keyword recognition convergence condition reached can also be that the target loss information is less than a specified threshold. In the embodiments of the present application, the preset training number and the specified threshold can be set in advance in combination with the training speed and accuracy of the network in actual application.

[0175] As can be seen from the above embodiments, by introducing the margin loss function of the face recognition model for network training, the network training result is better approximated to the ranking result of the key word extraction related feature similarity calculation, and the recognition accuracy and network generalization ability of the note keyword recognition network are improved.

[0176] In an optional embodiment, the text recognition network, the key semantic extraction network, and the keyword screening network can be three independent neural networks. Referring to Figure 10 , Figure 10 is another network structure diagram provided by the embodiments of the present application. Specifically, based on the network structure as shown in Figure 10 , the note keyword recognition method based on note image provided by the embodiments of the present application can further include: collecting a note image of target note information; inputting the note image into the text recognition network for text recognition to obtain note recognition text corresponding to the note image; inputting the note recognition text into the key semantic extraction network for key semantic extraction to obtain target semantic feature information; obtaining a plurality of candidate keywords corresponding to the note recognition text; inputting the target semantic feature information and the plurality of candidate keywords into the keyword screening network for keyword screening to determine the note keyword in the plurality of candidate keywords.

[0177] Specifically, since the text recognition network, the key semantic extraction network and the keyword screening network are independent, the three networks can be trained respectively, and then the three trained networks are combined to obtain the trained note keyword recognition network. Alternatively, the three networks can be jointly trained to obtain the trained note keyword recognition network.

[0178] As can be seen from the technical solutions provided by the embodiments of the present application, in the application scenario of note keyword recognition of a note image, the text recognition network and the keyword screening network based on the shared context semantic recognition layer are used to generate an end-to-end note keyword recognition network, which can reduce the data size of the network model. By inputting the collected note image containing target note information into the text recognition network in the note keyword recognition network for text recognition, the note recognition text corresponding to the note image is obtained. Then, the note recognition text is input into the key semantic extraction network in the note keyword recognition network for key semantic extraction, and the target semantic feature information is obtained. Finally, the target semantic feature information and the multiple candidate keywords corresponding to the note recognition text are input into the keyword screening network in the note keyword recognition network for keyword screening to determine the note keyword in the multiple candidate keywords. This not only enables the note content information of the picture format note to be quickly recognized, but also enables the note content information to be quickly parsed and key semantics to be extracted, thereby improving the efficiency and accuracy of note keyword recognition of the note image, solving the problem that the current note recording tool cannot fully parse the key content of the picture format note, and greatly improving the use effect and user experience of the note recording tool.

[0179] The embodiments of the present application also provide a note recording device, which can include an image acquisition module and a processor, wherein:

[0180] The image acquisition module is configured to acquire a note image of target note information.

[0181] The processor is configured to perform text recognition on the note image input into the text recognition network in the note keyword recognition network to obtain note recognition text corresponding to the note image; perform key semantic extraction on the note recognition text input into the key semantic extraction network in the note keyword recognition network to obtain target semantic feature information; obtain multiple candidate keywords corresponding to the note recognition text; perform keyword screening on the target semantic feature information and the multiple candidate keywords input into the keyword screening network in the note keyword recognition network to determine a note keyword in the multiple candidate keywords; and the text recognition network and the keyword screening network share the context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on feature information.

[0182] In a specific embodiment, the image acquisition module can include a camera, and correspondingly, the note image of the target note information can include: capturing the target note information based on the camera to obtain the note image.

[0183] In a specific embodiment, the note keyword recognition network can be an end-to-end network model including a text recognition network, a key semantic extraction network, and a keyword screening network, wherein the text recognition network and the keyword screening network share a context semantic recognition layer.

[0184] In a specific embodiment, the trained end-to-end note keyword recognition network can be pre-downloaded from the cloud and built into the note recording device. After the image acquisition module of the note recording device acquires the note image of the target note information in real time, the processor of the note recording device can perform real-time note keyword recognition on the note image based on the built-in note keyword recognition network to obtain the note keyword.

[0185] In an optional embodiment, after the image acquisition module of the note recording device acquires the note image of the target note information in real time, the processor of the note recording device can upload the note image to the cloud, and the cloud can perform note keyword recognition on the note image based on the trained note keyword recognition network to obtain the note keyword.

[0186] Specifically, the detailed content of the note keyword recognition network in the embodiment of the note recording device can refer to the related detailed content of the note keyword recognition network in the embodiment of the note keyword recognition method based on the note image, which will not be repeated here.

[0187] Taking the application scenario of teaching and listening as an example, during the class, the note recording device provided in the embodiment of the present application can capture the PPT courseware played by the teacher to obtain a note picture, convert the note picture into a note recognition text through text recognition, further extract the key information of the note recognition text, recognize the note key semantic information, complete the note recording, and greatly improve the effect and user experience of note recording, solving the scene pain point.

[0188] In addition to this, applications related to picture and other format text note recording all belong to the potential application scenarios of the invention.

[0189] It should be noted that the device in the device embodiment and the method embodiment are based on the same inventive concept.

[0190] The embodiment of the present application also provides a note keyword recognition device based on a note image, as shown in Figure 11 The note keyword recognition device based on a note image can include:

[0191] The image acquisition module 1110 is configured to acquire a note image of target note information.

[0192] The text recognition module 1120 is configured to input the note image into a text recognition network in the note keyword recognition network to perform text recognition, to obtain note recognition text corresponding to the note image.

[0193] The key semantic extraction module 1130 is configured to input the note recognition text into a key semantic extraction network in the note keyword recognition network to perform key semantic extraction, to obtain target semantic feature information.

[0194] The candidate keyword acquisition module 1140 is configured to acquire a plurality of candidate keywords corresponding to the note recognition text.

[0195] The keyword screening module 1150 is configured to input the target semantic feature information and the plurality of candidate keywords into a keyword screening network in the note keyword recognition network to perform keyword screening, to determine a note keyword in the plurality of candidate keywords.

[0196] The text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on the feature information.

[0197] In the embodiments of the present disclosure, the image acquisition module 1110 described above can include:

[0198] The to-be-processed note image acquisition unit is configured to acquire a to-be-processed note image of target note information.

[0199] The grayscale processing unit is configured to perform grayscale processing on the to-be-processed note image, to obtain a first note image.

[0200] The binarization processing unit is configured to perform binarization processing on the first note image, to obtain a second note image.

[0201] The inclination correction processing unit is configured to perform inclination correction processing on the second note image, to obtain a third note image.

[0202] The noise filtering processing unit is configured to perform noise filtering processing on the third note image, to obtain a fourth note image.

[0203] The normalization processing unit is configured to perform normalization processing on the fourth note image, to obtain the note image.

[0204] In one specific embodiment, the text recognition network can further include an image feature extraction layer and a text transcription layer, and the text recognition module 1120 described above can include:

[0205] The image feature extraction unit is configured to input the note image into an image feature extraction layer to perform image feature extraction, and obtain an image feature sequence.

[0206] The context semantic recognition unit is configured to input the image feature sequence into a context semantic recognition layer to perform context semantic recognition, and obtain character feature information corresponding to the image feature sequence.

[0207] The feature conversion unit is configured to input the character feature information into a text transcription layer to perform feature conversion, and obtain the note recognition text.

[0208] In an embodiment, the key semantic extraction network includes a topic extraction network, a high-frequency semantic mining network, and a semantic fusion network. The key semantic extraction module 1130 includes:

[0209] The topic extraction unit is configured to input the note recognition text into the topic extraction network to perform topic extraction, and obtain topic word feature information.

[0210] The high-frequency semantic mining unit is configured to input the note recognition text into the high-frequency semantic mining network to perform high-frequency semantic mining, and obtain high-frequency semantic feature information.

[0211] The semantic fusion unit is configured to input the topic word feature information and the high-frequency semantic feature information into the semantic fusion network to perform semantic fusion, and obtain target semantic feature information.

[0212] In an embodiment, the high-frequency semantic feature information includes at least one high-frequency keyword feature information and a semantic support degree corresponding to each high-frequency keyword feature information. The semantic fusion network includes a weighting layer and a fusion layer. The semantic fusion unit includes:

[0213] The weighting processing unit is configured to input the semantic support degree and the at least one high-frequency keyword feature information into the weighting layer, perform weighting processing on the at least one high-frequency keyword feature information based on the semantic support degree, and obtain initial semantic feature information.

[0214] The feature fusion unit is configured to input the initial semantic feature information and the topic word feature information into the fusion layer to perform semantic fusion, and obtain the target semantic feature information.

[0215] In an embodiment, the keyword screening network further includes an association analysis layer and a keyword screening layer. The keyword screening module 1150 includes:

[0216] The context semantic recognition unit is configured to input the target semantic feature information into a context semantic recognition layer to perform context semantic recognition, and obtain context semantic features.

[0217] The feature correlation analysis unit is configured to perform feature correlation analysis on the context semantic feature and the plurality of candidate keywords in the feature correlation analysis layer to obtain feature correlation information of each candidate keyword and the context semantic feature.

[0218] The keyword screening unit is configured to perform keyword screening on the feature correlation information in the keyword screening layer to determine the note keyword.

[0219] In the embodiments of the present disclosure, the note keyword recognition network is trained by the following device:

[0220] The sample acquisition module is configured to acquire a sample note image and a labeled note keyword corresponding to the sample note image.

[0221] The sample text recognition module is configured to input the sample note image into a preset text recognition network in the preset note keyword recognition network to perform text recognition to obtain sample note recognition text corresponding to the sample note image.

[0222] The sample key semantic extraction module is configured to input the sample note recognition text into a preset key semantic extraction network in the preset note keyword recognition network to perform key semantic extraction to obtain sample semantic feature information.

[0223] The sample candidate keyword acquisition module is configured to acquire a plurality of sample candidate keywords corresponding to the sample note recognition text.

[0224] The sample keyword screening module is configured to input the sample semantic feature information and the plurality of sample candidate keywords into a preset keyword screening network in the preset note keyword recognition network to perform keyword screening to determine a sample note keyword in the plurality of sample candidate keywords.

[0225] The target loss information determination module is configured to determine target loss information based on the labeled note keyword and the sample note keyword.

[0226] The network training module is configured to train the preset note keyword recognition network based on the target loss information to generate the note keyword recognition network.

[0227] The preset text recognition network and the preset keyword screening network share a preset context semantic recognition layer, and the preset context semantic recognition layer is configured to perform context semantic recognition on the sample feature information.

[0228] It should be noted that the device in the device embodiment and the method embodiment are based on the same inventive concept.

[0229] The embodiment of the present application provides a note keyword recognition device based on a note image, the note keyword recognition device based on the note image comprises a processor and a memory, the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to realize the note keyword recognition method based on the note image provided in the above method embodiment.

[0230] Further, Figure 12 A hardware structure schematic diagram of a note keyword recognition device based on a note image for realizing the note keyword recognition method based on the note image provided in the embodiment of the present application is shown, and the note keyword recognition device based on the note image can participate in constituting or containing the note keyword recognition apparatus based on the note image provided in the embodiment of the present application. As shown in the figure, Figure 12 The note keyword recognition device based on the note image 120 can comprise one or more (processors 1202a, 1202b,..., 1202n are shown in the figure) processors 1202 (the processor 1202 can comprise but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1204 for storing data, and a transmission device 1206 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those skilled in the art can understand that, Figure 12 The structure shown in the figure is only schematic, and it does not limit the structure of the above-mentioned electronic device. For example, the note keyword recognition device based on the note image 120 can further comprise more or less components than those shown in the figure, or have a different configuration from that shown in the figure. Figure 12 For example, the note keyword recognition device based on the note image 120 can further comprise more or less components than those shown in the figure, or have a different configuration from that shown in the figure. Figure 12 For example, the note keyword recognition device based on the note image 120 can further comprise more or less components than those shown in the figure, or have a different configuration from that shown in the figure.

[0231] It should be noted that the one or more processors 1202 and / or other data processing circuits described above can be referred to as "data processing circuits" herein. The data processing circuit can be embodied in whole or in part as software, hardware, firmware or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or any one of the other elements combined into the note keyword recognition device based on the note image 120 (or a mobile device) in whole or in part. As referred to in the embodiment of the present application, the data processing circuit controls as a kind of processor (for example, the selection of the variable resistance terminal path connected with the interface).

[0232] The memory 1204 can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the note keyword recognition method based on note images as described in embodiments of the present application. The processor 1202 executes various functional applications and data processing by running the software programs and modules stored in the memory 1204, i.e., implements the note keyword recognition method based on note images as described above. The memory 1204 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 1204 can further include memories disposed remotely with respect to the processor 1202, which can be connected to the note keyword recognition device 120 based on note images through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0233] The transmission device 1206 is configured to receive or send data via a network. Examples of the network include a wireless network provided by a communication provider of the note keyword recognition device 120 based on note images. In one example, the transmission device 1206 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one embodiment, the transmission device 1206 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.

[0234] The display can be, for example, a touch screen type liquid crystal display (LCD) that enables a user to interact with a user interface of the note keyword recognition device 120 (or mobile device) based on note images.

[0235] Embodiments of the present application also provide a computer readable storage medium that can be disposed in the note keyword recognition device based on note images to save at least one instruction or at least one program for implementing the note keyword recognition method based on note images in the method embodiments. The at least one instruction or the at least one program is loaded and executed by the processor to implement the note keyword recognition method based on note images provided by the method embodiments.

[0236] Optionally, in the embodiment, the storage medium can be located in at least one of the plurality of network servers of the computer network. Optionally, in the embodiment, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media capable of storing program codes.

[0237] Embodiments of the present application also provide a computer program product or computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the note keyword identification method based on a note image as provided in the method embodiments.

[0238] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of the present application are described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that in the embodiments and still achieve the desired result. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multi-task processing and parallel processing are possible or advantageous.

[0239] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts of each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0240] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.

[0241] The above is only the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A note keyword recognition method based on a note image, characterized by, The method comprises: collecting a note image of target note information; inputting the note image into a text recognition network in a note keyword recognition network for text recognition to obtain note recognition text corresponding to the note image; inputting the note recognition text into a key semantic extraction network in the note keyword recognition network for key semantic extraction to obtain target semantic feature information; obtaining a plurality of candidate keywords corresponding to the note recognition text; inputting the target semantic feature information and the plurality of candidate keywords into a keyword screening network in the note keyword recognition network for keyword screening to determine a note keyword in the plurality of candidate keywords; wherein the text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on feature information; The text recognition network further comprises an image feature extraction layer and a text transcription layer, and the inputting the note image into a text recognition network in a note keyword recognition network for text recognition to obtain note recognition text corresponding to the note image comprises: inputting the note image into the image feature extraction layer for image feature extraction to obtain an image feature sequence; inputting the image feature sequence into the context semantic recognition layer for context semantic recognition to obtain character feature information corresponding to the image feature sequence; inputting the character feature information into the text transcription layer for feature conversion to obtain the note recognition text.

2. The method of claim 1, wherein, The key semantic extraction network comprises a theme extraction network, a high-frequency semantic mining network and a semantic fusion network, and the inputting the note recognition text into a key semantic extraction network in a note keyword recognition network for key semantic extraction to obtain target semantic feature information comprises: inputting the note recognition text into the theme extraction network for theme extraction to obtain theme keyword feature information; inputting the note recognition text into the high-frequency semantic mining network for high-frequency semantic mining to obtain high-frequency semantic feature information; inputting the theme keyword feature information and the high-frequency semantic feature information into the semantic fusion network for semantic fusion to obtain target semantic feature information.

3. The method of claim 2, wherein, The high-frequency semantic feature information comprises at least one high-frequency keyword feature information and a semantic support degree corresponding to each of the at least one high-frequency keyword feature information, the semantic fusion network comprises a weighting layer and a fusion layer, and the inputting the theme keyword feature information and the high-frequency semantic feature information into the semantic fusion network for semantic fusion to obtain target semantic feature information comprises: inputting the semantic support degree and the at least one high-frequency keyword feature information into the weighting layer, performing weighted processing on the at least one high-frequency keyword feature information based on the semantic support degree to obtain initial semantic feature information; inputting the initial semantic feature information and the theme keyword feature information into the fusion layer for semantic fusion to obtain the target semantic feature information.

4. The method of claim 1, wherein, The keyword screening network further comprises an association analysis layer and a keyword screening layer, and the keyword screening network in the note keyword recognition network is used for keyword screening on the basis of the target semantic feature information and the plurality of candidate keywords, so as to determine the note keyword in the plurality of candidate keywords. The target semantic feature information is input into the context semantic recognition layer for context semantic recognition, so as to obtain a context semantic feature; The context semantic feature and the plurality of candidate keywords are input into the association analysis layer for feature association analysis, so as to obtain feature association information of each candidate keyword and the context semantic feature; The feature association information is input into the keyword screening layer for keyword screening, so as to determine the note keyword.

5. The method according to any one of claims 1 to 4, characterized in that, The note image of the target note information comprises: A to-be-processed note image of the target note information is collected; The to-be-processed note image is subjected to grayscale processing, so as to obtain a first note image; The first note image is subjected to binarization processing, so as to obtain a second note image; The second note image is subjected to tilt correction processing, so as to obtain a third note image; The third note image is subjected to noise filtering processing, so as to obtain a fourth note image; The fourth note image is subjected to normalization processing, so as to obtain the note image.

6. The method according to any one of claims 1 to 4, characterized in that, The note keyword recognition network is trained in the following manner: A sample note image and a labeled note keyword corresponding to the sample note image are obtained; The sample note image is input into a preset text recognition network in a preset note keyword recognition network for text recognition, so as to obtain a sample note recognition text corresponding to the sample note image; The sample note recognition text is input into a preset key semantic extraction network in the preset note keyword recognition network for key semantic extraction, so as to obtain sample semantic feature information; A plurality of sample candidate keywords corresponding to the sample note recognition text are obtained; The sample semantic feature information and the plurality of sample candidate keywords are input into a preset keyword screening network in the preset note keyword recognition network for keyword screening, so as to determine a sample note keyword in the plurality of sample candidate keywords; On the basis of the labeled note keyword and the sample note keyword, target loss information is determined; On the basis of the target loss information, the preset note keyword recognition network is trained, so as to generate the note keyword recognition network; The preset text recognition network and the preset keyword screening network share a preset context semantic recognition layer, and the preset context semantic recognition layer is used for context semantic recognition on sample feature information.

7. A pen recording apparatus characterized by comprising: The note input device comprises an image collection module and a processor, and wherein: The image collection module is used for collecting a note image of target note information; The processor is configured to perform text recognition on the note image in a text recognition network in the note keyword recognition network to obtain note recognition text corresponding to the note image, perform keyword semantic extraction on the note recognition text in a keyword semantic extraction network in the note keyword recognition network to obtain target semantic feature information, obtain a plurality of candidate keywords corresponding to the note recognition text, and perform keyword screening on the target semantic feature information and the plurality of candidate keywords in a keyword screening network in the note keyword recognition network to determine a note keyword in the plurality of candidate keywords. The text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on feature information. The text recognition network further includes an image feature extraction layer and a text transcription layer. The processor is further configured to input the note image into the image feature extraction layer to perform image feature extraction to obtain an image feature sequence, input the image feature sequence into the context semantic recognition layer to perform context semantic recognition to obtain character feature information corresponding to the image feature sequence, and input the character feature information into the text transcription layer to perform feature conversion to obtain the note recognition text.

8. The apparatus of claim 7, wherein, The keyword semantic extraction network includes a theme extraction network, a high-frequency semantic mining network, and a semantic fusion network. The processor is further configured to input the note recognition text into the theme extraction network to perform theme extraction to obtain theme keyword feature information, input the note recognition text into the high-frequency semantic mining network to perform high-frequency semantic mining to obtain high-frequency semantic feature information, and input the theme keyword feature information and the high-frequency semantic feature information into the semantic fusion network to perform semantic fusion to obtain the target semantic feature information.

9. The apparatus of claim 8, wherein, The high-frequency semantic feature information includes at least one high-frequency keyword feature information and semantic support degrees corresponding to the at least one high-frequency keyword feature information. The semantic fusion network includes a weighting layer and a fusion layer. The processor is further configured to input the semantic support degrees and the at least one high-frequency keyword feature information into the weighting layer, perform weighting processing on the at least one high-frequency keyword feature information based on the semantic support degrees to obtain initial semantic feature information, and input the initial semantic feature information and the theme keyword feature information into the fusion layer to perform semantic fusion to obtain the target semantic feature information.

10. The apparatus of claim 7, wherein, The keyword screening network further includes an association analysis layer and a keyword screening layer. The processor is further configured to input the target semantic feature information into the context semantic recognition layer to perform context semantic recognition to obtain context semantic features, input the context semantic features and the plurality of candidate keywords into the association analysis layer to perform feature association analysis to obtain feature association information of each candidate keyword and the context semantic features, and input the feature association information into the keyword screening layer to perform keyword screening to determine the note keyword.

11. The apparatus of any one of claims 7 to 10, wherein, The processor is further configured to: collect a to-be-processed note image of the target note information; perform grayscale processing on the to-be-processed note image to obtain a first note image; perform binaryzation processing on the first note image to obtain a second note image; perform tilt correction processing on the second note image to obtain a third note image; perform noise filtering processing on the third note image to obtain a fourth note image; perform normalization processing on the fourth note image to obtain the note image.

12. The apparatus of any one of claims 7 to 10, wherein, The processor is further configured to: obtain a sample note image and a labeled note keyword corresponding to the sample note image; input the sample note image into a preset text recognition network in a preset note keyword recognition network to perform text recognition, to obtain a sample note recognition text corresponding to the sample note image; input the sample note recognition text into a preset key semantic extraction network in the preset note keyword recognition network to perform key semantic extraction, to obtain sample semantic feature information; obtain a plurality of sample candidate keywords corresponding to the sample note recognition text; input the sample semantic feature information and the plurality of sample candidate keywords into a preset keyword screening network in the preset note keyword recognition network to perform keyword screening, to determine a sample note keyword in the plurality of sample candidate keywords; determine target loss information based on the labeled note keyword and the sample note keyword; train the preset note keyword recognition network based on the target loss information, to generate the note keyword recognition network; wherein the preset text recognition network and the preset keyword screening network share a preset context semantic recognition layer, and the preset context semantic recognition layer is configured to perform context semantic recognition on sample feature information.

13. A note keyword recognition apparatus based on a note image, characterized by, The apparatus comprises: an image collection module configured to collect a note image of target note information; a text recognition module configured to input the note image into a text recognition network in a note keyword recognition network to perform text recognition, to obtain note recognition text corresponding to the note image; a key semantic extraction module configured to input the note recognition text into a key semantic extraction network in the note keyword recognition network to perform key semantic extraction, to obtain target semantic feature information; a candidate keyword obtaining module configured to obtain a plurality of candidate keywords corresponding to the note recognition text; a keyword screening module configured to input the target semantic feature information and the plurality of candidate keywords into a keyword screening network in the note keyword recognition network to perform keyword screening, to determine a note keyword in the plurality of candidate keywords; wherein the text recognition network and the keyword screening network share a context semantic recognition layer, and the context semantic recognition layer is configured to perform context semantic recognition on feature information; the text recognition network further comprises an image feature extraction layer and a text transcription layer, and the text recognition module comprises: an image feature extraction unit configured to input the note image into the image feature extraction layer to perform image feature extraction, to obtain an image feature sequence; The context semantic recognition unit is configured to input the image feature sequence into the context semantic recognition layer for context semantic recognition, so as to obtain character feature information corresponding to the image feature sequence. The feature conversion unit is configured to input the character feature information into the text transcription layer for feature conversion, so as to obtain the note recognition text.

14. The apparatus of claim 13, wherein, The key semantic extraction network comprises a theme extraction network, a high-frequency semantic mining network and a semantic fusion network, and the key semantic extraction module comprises: The theme extraction unit is configured to input the note recognition text into the theme extraction network for theme extraction, so as to obtain theme word feature information. The high-frequency semantic mining unit is configured to input the note recognition text into the high-frequency semantic mining network for high-frequency semantic mining, so as to obtain high-frequency semantic feature information. The semantic fusion unit is configured to input the theme word feature information and the high-frequency semantic feature information into the semantic fusion network for semantic fusion, so as to obtain target semantic feature information.

15. The apparatus of claim 14, wherein, The high-frequency semantic feature information comprises at least one high-frequency keyword feature information and semantic support degrees corresponding to the at least one high-frequency keyword feature information, the semantic fusion network comprises a weighting layer and a fusion layer, and the semantic fusion unit comprises: The weighting processing unit is configured to input the semantic support degrees and the at least one high-frequency keyword feature information into the weighting layer, perform weighting processing on the at least one high-frequency keyword feature information based on the semantic support degrees, and obtain initial semantic feature information. The feature fusion unit is configured to input the initial semantic feature information and the theme word feature information into the fusion layer for semantic fusion, so as to obtain the target semantic feature information.

16. The apparatus of claim 13, wherein, The keyword screening network further comprises an association analysis layer and a keyword screening layer, and the keyword screening module comprises: The context semantic recognition unit is configured to input the target semantic feature information into the context semantic recognition layer for context semantic recognition, so as to obtain context semantic features. The feature association analysis unit is configured to input the context semantic features and the plurality of candidate keywords into the association analysis layer for feature association analysis, so as to obtain feature association information of each candidate keyword and the context semantic features. The keyword screening unit is configured to input the feature association information into the keyword screening layer for keyword screening, so as to determine the note keyword.

17. The apparatus of any one of claims 13 to 16, wherein, The image acquisition module comprises: The to-be-processed note image acquisition unit is configured to acquire a to-be-processed note image of the target note information. The grayscale processing unit is configured to perform grayscale processing on the to-be-processed note image, so as to obtain a first note image. The binarization processing unit is configured to perform binarization processing on the first note image, so as to obtain a second note image. The inclination correction processing unit is configured to perform inclination correction processing on the second note image, so as to obtain a third note image. The noise filtering processing unit is configured to perform noise filtering processing on the third note image, so as to obtain a fourth note image. The normalization processing unit is configured to perform normalization processing on the fourth note image, so as to obtain the note image.

18. The apparatus of any one of claims 13 to 16, wherein, The note keyword recognition network is trained by an apparatus including: a sample obtaining module configured to obtain a sample note image and a labeled note keyword corresponding to the sample note image; a sample text recognition module configured to input the sample note image into a preset text recognition network in a preset note keyword recognition network to perform text recognition to obtain a sample note recognition text corresponding to the sample note image; a sample key semantic extraction module configured to input the sample note recognition text into a preset key semantic extraction network in the preset note keyword recognition network to perform key semantic extraction to obtain sample semantic feature information; a sample candidate keyword obtaining module configured to obtain a plurality of sample candidate keywords corresponding to the sample note recognition text; a sample keyword screening module configured to input the sample semantic feature information and the plurality of sample candidate keywords into a preset keyword screening network in the preset note keyword recognition network to perform keyword screening to determine a sample note keyword in the plurality of sample candidate keywords; a target loss information determining module configured to determine target loss information based on the labeled note keyword and the sample note keyword; a network training module configured to train the preset note keyword recognition network based on the target loss information to generate the note keyword recognition network. The preset text recognition network and the preset keyword screening network share a preset context semantic recognition layer, and the preset context semantic recognition layer is configured to perform context semantic recognition on sample feature information.

19. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the note keyword recognition method based on a note image according to any one of claims 1 to 6.

20. A computer program product, characterised in that, The computer program product includes at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the note keyword recognition method based on a note image according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Case record processing method and device, equipment and medium

    CN109800304A

  • Voice synthesis method and device

    CN111968618A