Named entity recognition method and device, equipment, medium and program product
By acquiring multimodal data of financial text and images, and using image feature vectors and large language models to identify financial terms, the problem of text-image separation is solved, and the accuracy and robustness of financial named entity recognition are improved.
Patent Information
- Application Number
- CN202511052372.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional named entity recognition methods rely solely on text data, making it difficult to accurately identify financial entities where text and images are closely intertwined in financial scenarios.
Acquire multimodal financial data, including financial text and images, identify financial terms through financial image feature vectors and large language models, concatenate them with text word vectors to form a vector sequence, and input it into a named entity model for recognition.
It significantly improves the accuracy and robustness of financial named entity recognition and reduces the risk of mislabeling due to missing text information.
Smart Images

Figure CN120911463A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a named entity recognition method, device, equipment, medium and program product. BACKGROUND
[0002] In the financial field, named entity recognition (NER) technology is a key foundation for information extraction, which accurately identifies core entities such as financial institution names, financial product names, securities codes, amounts, dates, etc. in massive financial texts, providing indispensable data support for financial information structured processing, risk assessment model construction, and investment decision support. With the explosive growth of financial data, efficient and accurate NER technology has become a core link in intelligent investment research, risk monitoring, and other financial technology applications.
[0003] Existing NER technology mainly relies on three methods: 1. Rule method: based on manually defined regular expressions or dictionary matching, suitable for structured data (such as fixed format reports), but difficult to generalize to complex and variable financial texts; 2. Statistical model method: such as Hidden Markov Model (HMM) and Maximum Entropy Model, trained to improve generalization ability through labeled data, but heavily dependent on feature engineering, with limited adaptability to financial professional terms; 2. Deep learning method: based on RNN, LSTM, GRU combined with CRF model, which can capture text context information and significantly improve accuracy, but still has bottlenecks in handling long-distance dependencies (such as cross-paragraph entity association).
[0004] However, traditional named entity recognition methods only rely on text data, and in the financial scenario, when the text is closely related to the image (such as financial statement charts and loan application text), due to the reliance on text data only, it cannot accurately identify entities. SUMMARY
[0005] The present application provides a named entity recognition method, device, equipment, medium and program product to solve the technical problem that the traditional named entity recognition (NER) system only relies on text data and cannot accurately identify named entities when the text is closely related to the image in the financial scenario.
[0006] In a first aspect, the present application provides a named entity recognition method, comprising:
[0007] Obtaining multi-modal financial data of named entities to be recognized, the multi-modal financial data comprising: a financial text, and a financial image associated with the financial text;
[0008] Based on the financial image, obtaining a financial image feature vector;
[0009] identify the financial nouns in the financial image based on the financial image feature vector and financial knowledge by using a large language model;
[0010] convert words in the financial text into word vectors;
[0011] concatenate the financial nouns and the word vectors to obtain a vector sequence;
[0012] identify the named entities in the financial text based on the vector sequence by using a named entity model.
[0013] In a second aspect, the present application provides a named entity identification device, the device comprising:
[0014] an acquisition module configured to acquire multi-modal financial data to be identified named entities, the multi-modal financial data comprising: a financial text and a financial image associated with the financial text;
[0015] a feature vector identification module configured to acquire a financial image feature vector based on the financial image;
[0016] a financial noun identification module configured to identify financial nouns in the financial image based on the financial image feature vector and financial knowledge by using a large language model;
[0017] a word vector conversion module configured to convert words in the financial text into word vectors;
[0018] a concatenation module configured to concatenate the financial nouns and the word vectors to obtain a vector sequence;
[0019] a named entity identification module configured to identify the named entities in the financial text based on the vector sequence by using a named entity model.
[0020] In a third aspect, the present application provides an electronic device comprising: a processor and a memory in communication with the processor;
[0021] the memory stores computer execution instructions;
[0022] the processor executes the computer execution instructions stored in the memory to implement the method of the first aspect.
[0023] In a fourth aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing computer execution instructions, the computer execution instructions being executed by a processor to implement the method of the first aspect.
[0024] In a fifth aspect, a computer program product comprises a computer program, which, when executed by a processor, implements the method of the first aspect.
[0025] The naming entity recognition method, device, equipment, medium and program product provided by the application achieve the following technical effects:
[0026] The scheme obtains multi-modal data including financial text and associated financial images, extracts a feature vector based on the images, and then uses a large language model to recognize financial nouns, and then splices these nouns and text word vectors into a sequence, and finally inputs a naming entity model for recognition, thereby making up for the lack of information in pure text analysis, solving the problem of image and text fragmentation, and thus significantly improving the accuracy and robustness of financial named entity recognition, while reducing the risk of incorrect labeling caused by missing text information. BRIEF DESCRIPTION OF DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application together with the description.
[0028] Figure 1 The naming entity recognition method flowchart provided by the embodiment of the application;
[0029] Figure 2 The naming entity recognition method implementation flowchart provided by the embodiment of the application;
[0030] Figure 3 The naming entity recognition device structure schematic diagram provided by the embodiment of the application;
[0031] Figure 4 The electronic device structure schematic diagram provided by the embodiment of the application.
[0032] Through the above drawings, the specific embodiments of the application have been shown, and will be described in more detail in the following. These drawings and written descriptions are not intended to limit the scope of the concept of the application by any means, but to illustrate the concept of the application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0033] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, the same numbers are used to indicate the same or similar components. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the application. Rather, they are merely examples of devices and methods consistent with some aspects of the application, as detailed in the appended claims.
[0034] It should be noted that the named entity recognition method, device, equipment, medium and program product provided by the present application can be used in the field of artificial intelligence, and can also be used in any field other than artificial intelligence. The application of the named entity recognition method, device, equipment, medium and program product in the present application is not limited.
[0035] In the financial field, named entity recognition (NER) technology plays a crucial role in accurately identifying key entities such as financial institution names, financial product names, security codes, amounts, dates, etc. from massive financial texts, providing important data support for financial information extraction, risk assessment, investment decision-making, etc. Traditional NER methods mainly rely on rules, statistical models, and deep learning models based on recurrent neural networks (RNN), long short-term memory networks (LSTM), gated recurrent units (GRU) combined with conditional random fields (CRF). Rule-based methods use manually defined regular expressions or dictionary matching to identify specific entities, which are suitable for structured data but difficult to handle complex and variable financial texts. Statistical model-based methods (such as hidden Markov models and maximum entropy models) are trained using labeled data, which can improve generalization to some extent, but are highly dependent on feature engineering. Deep learning-based methods (such as RNN, LSTM, GRU combined with CRF) can capture the context information in the text, significantly improving the accuracy of NER, but still have limitations in handling long-distance dependencies and multi-modal data.
[0036] Financial texts often contain a large number of complex professional terms and are closely related to financial images (such as charts in financial reports, transaction record screenshots, financial bills, etc.). When the text involves financial entities that need to be understood in combination with images, simple text-based analysis methods cannot accurately identify all entities.
[0037] With the rapid development of large language models (LLM), their powerful language understanding and generation capabilities provide new solutions for financial NER tasks. Since LLM can learn language knowledge from massive text data combined with image information extracted by computer vision technology, it can solve the technical problem that simple text-based analysis methods cannot accurately identify all entities. Therefore, the technical concept of the present application is: inputting the multi-modal financial data to be identified, including financial text and associated financial images, processing the financial images to obtain their feature vectors. Then, based on the feature vectors of the financial images and financial knowledge, the large language model is used to identify financial nouns in the financial system. After converting the identified financial nouns into word vectors, the financial nouns and word vectors are concatenated into a vector sequence, and then a named entity model is used to identify named entities in the financial text based on the vector sequence.
[0038] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes can not be described again in some examples. The embodiments of the present application will be described below with reference to the drawings.
[0039] Figure 1 A named entity recognition method flow diagram is provided for the embodiments of the present application, and the method can be applied to a server. As shown in the figure, Figure 1 The method comprises the following steps.
[0040] S101, acquiring multi-modal financial data of a named entity to be recognized, the multi-modal financial data comprising: financial text, and a financial image associated with the financial text;
[0041] In financial business, text and image often corroborate each other, for example, loan text needs to be verified by financial statement images, so text and image data of two modalities need to be collected at the same time. Specifically, the financial text can be obtained from OCR recognition results, database fields or API interfaces, for example, enterprise loan application text of a bank credit system; the associated financial image refers to visual data semantically associated with the text, for example, balance sheet image, transaction flow screenshot, etc. with the same business ID matched through file metadata.
[0042] S102, acquiring a financial image feature vector based on the financial image;
[0043] In this step, for example, a region detection algorithm (such as a watershed algorithm or a deep target detection model YOLO) can be used to locate a key data region, and a sub-image such as an "asset amount area" in a financial statement can be segmented. Further, a pre-trained convolutional network such as the convolutional layer of ResNet or VGG can be used to extract the financial image feature vector.
[0044] S103, identifying a financial noun in the financial image based on the financial image feature vector and financial knowledge using a large language model;
[0045] In this step, the knowledge in the financial field is used to guide the large language model to identify the financial noun in the financial image. Specifically, for example, the financial image feature can be input into an MLP encoder, and the financial knowledge base (such as a term table of "short-term loans" and "fixed assets") can be one-hot encoded, and the two can be concatenated into a fusion vector. Further, the fusion vector and the business prompt text (for example, "identify the financial noun in the balance sheet") are input into a large language model (such as an LLM model) to generate a financial noun. The financial noun can be, for example, "fixed assets: 1000.00 million yuan".
[0046] S104, converting words in the financial text into word vectors;
[0047] Specifically, for example, the Word2Vec model (gensim.models.Word2Vec, window=5, min_word_freq=5) can be used to convert the words in the financial text into 300-dimensional word vectors.
[0048] S105, splice the financial noun and the word vector to obtain a vector sequence;
[0049] The purpose of this step is to align the text and image features by splicing, so that the NER model can perform collaborative analysis. Specifically, the financial noun vector and the text word vector can be spliced along the feature dimension, for example, the noun vector is copied to each position of the text sequence to generate a fusion sequence of [word vector, noun vector].
[0050] S106, using a named entity model, identifying the named entity in the financial text based on the vector sequence.
[0051] In this step, the vector sequence can be input into the BERT-CRF model to output entity labels, which can be, for example: "300 million yuan" - amount, "the enterprise" - organization name.
[0052] The named entity recognition method provided by the embodiment has the following technical effects:
[0053] The method obtains multi-modal data including financial text and associated financial images, extracts feature vectors based on the images, uses a large language model to identify financial nouns, then splices these nouns with text word vectors into a sequence, and finally inputs the named entity model for recognition, thereby making up for the lack of information in pure text analysis, solving the problem of image and text fragmentation, and significantly improving the accuracy and robustness of financial named entity recognition, while reducing the risk of false labeling caused by missing text information
[0054] Figure 2 The named entity recognition method provided by the embodiment of the present application is implemented as a flowchart, and the implementation process of the named entity recognition method provided by the embodiment of the present application will be described in detail. Figure 2 The implementation process of the named entity recognition method provided by the embodiment of the present application will be described.
[0055] Optionally, based on the financial image, a financial image feature vector is obtained, including:
[0056] I. Image segmentation is performed on the financial image to obtain at least one sub-image containing financial key data;
[0057] This step corresponds to Figure 2the financial image preprocessing step in the method. Since a financial image (such as a balance sheet) contains multiple functional areas (for example, including an asset / liability area), the key area needs to be segmented to focus on the most critical core data, so the image segmentation technique is used to extract a sub-image containing financial key data.
[0058] Optionally, the financial image is subjected to image segmentation, including:
[0059] converting the financial image into a grayscale image;
[0060] performing noise reduction processing on the grayscale image;
[0061] performing image enhancement processing on the noise-reduced grayscale image;
[0062] performing image segmentation based on the image-enhanced grayscale image to obtain at least one sub-image containing financial key data.
[0063] Specifically, the image segmentation step can be implemented through the following process: first, convert the color image into a grayscale image to reduce computational complexity, then use Gaussian filtering (kernel size 5x5, sigma=1.5) to reduce noise and eliminate scanning noise, then enhance the contrast of the text through histogram equalization, and finally segment the asset item area, amount area, and other sub-images based on Canny edge detection (low threshold 50, high threshold 150). By converting the color image into a grayscale image, the subsequent image processing steps can be simplified, as the grayscale image has only one channel (rather than three channels for RGB images), thereby reducing computational complexity; noise reduction processing can remove random noise in the image, improving the quality and clarity of the image and preventing misjudgment or incorrect segmentation results that may be caused by noise; image enhancement can improve the contrast and details of the image, making the key features in the image more obvious.
[0064] Overall, the technical effect of this step is that the subsequent feature extraction is only performed on the effective area, thereby improving the processing efficiency.
[0065] II. Extracting a financial feature vector from the sub-image using an image feature extraction model, the financial feature vector including visual features of financial elements;
[0066] This step corresponds Figure 2The CNN feature extraction step in the financial image feature extraction. Specifically, the feature extraction step can be operated as follows: a specific CNN architecture, i.e., the convolutional layer of the VGG16 network, is used to process the sub-image, wherein the VGG16 model can be built using the Keras framework, the pre-trained model is loaded through the keras.applications.VGG16() function, and the 2048-dimensional feature vector is output using the model.predict(), which contains the visual patterns (such as edge direction and character spacing) of the numbers / letters in the chart. The technical effect of this step is that the financial feature vector extracted by the CNN architecture as a kind of deep feature can accurately represent the financial elements and provide high information density input for subsequent noun recognition.
[0067] III. Normalizing the financial feature vector to obtain a financial image feature vector.
[0068] It should be noted that the original financial feature vector extracted due to the difference in dimension will cause difficulty in model convergence, so it is necessary to normalize the numerical scale to enable comparison of features of different images. Specifically, the normalization step can use L2 normalization (call sklearn.preprocessing.normalize() function) to eliminate amplitude differences but preserve direction information.
[0069] Optionally, based on the financial image feature vector and the financial knowledge, a large language model is used to identify the financial nouns in the financial image, including:
[0070] I. The financial image feature vector and the financial knowledge are respectively encoded and processed.
[0071] This step corresponds to Figure 2 The feature encoding and fusion step in the "proper noun extraction". Since the financial image feature vector (such as the balance sheet visual feature) and the financial knowledge (such as the financial statement item terminology) belong to different modalities, they need to be encoded into a unified mathematical representation to support fusion, so a differentiated encoding strategy is adopted. Specifically, the encoding processing of the financial image feature vector can be realized by a multilayer perceptron, which uses a fully connected layer with a nonlinear activation function to map high-dimensional visual features to a semantic space; the encoding processing of the financial knowledge can be realized by using a one-hot encoding method to convert predefined financial institution names, financial item names, and other discrete labels into numerical vectors. The technical effect is that the visual features and domain knowledge are aligned in the same vector space, thereby supporting subsequent multi-modal fusion.
[0072] II. The encoded financial image feature vector and the encoded financial knowledge are spliced to obtain a fusion feature vector.
[0073] Specifically, the concatenation process can be performed as follows: The image encoding vector output by the multilayer perceptron is concatenated end-to-end with the knowledge vector generated by one-hot encoding at the final dimension, forming a fused feature vector containing both visual features and financial knowledge. The technical effect is to enhance feature representation capabilities, enabling large language models to simultaneously perceive image content and financial domain rules.
[0074] Third, input the fused feature vectors and the prompt text related to the business type corresponding to the financial image into the large model to obtain the financial terms in the financial image.
[0075] This step corresponds to Figure 2 The "Property Noun Extraction" step includes the prompting engineering optimization step and the generation step within the generation and filtering steps. It's important to note that because the large-scale language model needs a clear generation target and context, prompt text is introduced to guide the generation direction, thereby optimizing entity extraction accuracy through prompting engineering. Specifically, the input to the large-scale model step can be implemented as follows: Based on the financial image business type (such as a balance sheet), construct prompt text (e.g., "Identify asset items, liability items, and amounts"), and input it along with the fused feature vector into a GPT-type model to output structured financial terms (e.g., "Fixed Assets: X million yuan").
[0076] Optionally, corresponding to the filtering step in the generation and filtering steps, after identifying financial terms in the financial image, the method further includes:
[0077] 1. Use formatting rules to check the format of financial terms;
[0078] Because financial terms generated by large language models may contain formatting errors (such as missing monetary units), it is necessary to filter out erroneous expressions using formatting rules. Specifically, formatting checks can be implemented using regular expressions: for monetary terms, set rules for numbers plus units (e.g., requiring the numerical value and the unit "yuan / ten thousand yuan"), and for date terms, set rules for year-month-day formatting. The technical effect is to exclude erroneous expressions, ensuring that subsequent processing only retains formatted terms.
[0079] 2. Perform a validity check on the financial terms that have passed the format check;
[0080] Since even formatted terms may contain semantic errors (e.g., "stock code" appearing on a balance sheet), secondary verification using domain knowledge is necessary. Therefore, machine learning models can be used to determine entity validity. Specifically, the validity check can be performed as follows: using a support vector machine classifier (trained on a labeled financial terminology dataset), the input term text vectors are used to output validity probability values, retaining only terms with a probability threshold. The technical effect of this step is to identify and remove invalid terms irrelevant to the context, improving data reliability.
[0081] III. Retaining the financial terms that pass the validity check;
[0082] The financial terms that pass the validity check are spliced with the word vectors to obtain a vector sequence, which includes:
[0083] The financial terms that pass the validity check are spliced with the word vectors to obtain a vector sequence.
[0084] Since the text and image information need to be analyzed in coordination, the financial terms that pass the validity check need to be aligned with the text word vectors in a unified space. Specifically, the splicing step can be implemented as follows: The financial term feature vector is copied and extended to the length of the text sequence, and is spliced with the word vector generated by Word2Vec along the feature dimension (such as [word vector, term vector]).
[0085] Corresponding to Figure 2 The named entity annotation in the "output result" part of the method further includes:
[0086] Post-processing operation one: if there is a first named entity with multiple annotations in the identified named entity, remove the redundant annotations in the first named entity; and / or
[0087] Post-processing operation two: if there is a second named entity in the identified named entity whose annotation does not satisfy the named entity rules in the financial field, correct the annotation of the second named entity; and / or
[0088] Post-processing operation three: if there is a third named entity with an error annotation in the identified named entity, correct the annotation of the third named entity.
[0089] For post-processing operation one: since the named entity model may generate multiple overlapping annotations for the same entity (such as "short-term loans" being repeatedly labeled as financial project names), this will interfere with downstream analysis, so redundant annotations need to be removed to maintain entity uniqueness. Specifically, the step of removing redundant annotations can be implemented by traversing the annotation sequence: when multiple annotations of the same type are detected at the same text location, only the annotation with the highest confidence or the first appearing annotation is retained, thereby avoiding financial data distortion caused by overlapping entity annotations in loan review and other scenarios.
[0090] For post-processing operation two: since part of the entity annotations may violate the naming specifications in the financial field (such as labeling "accounts receivable" as an organization name), they need to be forcibly corrected according to the pre-defined rule library. Specifically, the step of correcting the annotation can be operated as follows: a financial entity rule dictionary is constructed (such as "repayment period" must belong to time class entities), and when the "second named entity" annotation type conflicts with the dictionary rule, it is automatically corrected to the correct type. Thus, the label type misplacement problem is solved.
[0091] For post-processing operation three: for entities that the model obviously misjudges (such as labeling the non-amount text "repayment period" as an amount), the error labeling needs to be corrected. Specifically, the correction error labeling step may, for example, use regular expression + keyword library double verification: when detecting that the "third named entity" meets the following conditions at the same time: 1, does not conform to the amount / date format rule; 2, contains preset keywords such as "period" and "ratio", correct it to the corresponding entity type (such as time class or ratio class), thereby eliminating the risk of false recognition of key entities (such as loan amount).
[0092] Figure 3 The structure diagram of the named entity recognition device provided by the embodiment of the present application is shown in Figure 3 The device 30 includes:
[0093] The acquisition module 301 is configured to acquire multi-modal financial data of a named entity to be recognized, the multi-modal financial data including: a financial text and a financial image associated with the financial text.
[0094] The feature vector identification module 302 is configured to acquire a financial image feature vector based on the financial image.
[0095] The financial noun identification module 303 is configured to identify a financial noun in the financial image based on the financial image feature vector and financial knowledge by using a large language model.
[0096] The word vector conversion module 304 is configured to convert a word in the financial text into a word vector.
[0097] The splicing module 305 is configured to splice the financial noun and the word vector to obtain a vector sequence.
[0098] The named entity recognition module 306 is configured to identify a named entity in the financial text based on the vector sequence by using a named entity model.
[0099] Figure 4 The structure diagram of the electronic device provided by the embodiment of the present application is shown in Figure 4 The device 40 includes at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. The processor 401, the memory 402 and the communication component 403 are connected through a bus 404.
[0100] In the specific implementation process, the at least one processor 401 executes the computer execution instructions stored in the memory 402, so that the at least one processor 401 executes the above-mentioned method.
[0101] The specific implementation process of the processor 401 can refer to the method embodiments described above, which have similar implementation principles and technical effects, and will not be described here again.
[0102] The application further provides a computer-readable storage medium, which stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method described above is implemented.
[0103] The application further provides a computer program product, which comprises a computer program, and when the computer program is executed by the processor, the method described above is implemented.
[0104] It should be noted that, for each of the above method embodiments, in order to simply describe, each is described as a combination of a series of actions, but those skilled in the art should know that the application is not limited by the order of the actions described, because according to the application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the application.
[0105] It should be further noted that, although each step in the flowchart is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise stated in this document, the execution of these steps has no strict order limitation, and these steps can be executed in other order. Moreover, at least part of the steps in the flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or sub-steps or stages of other steps.
[0106] It should be understood that the above-described device embodiments are only illustrative, and the device of the application can also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is only a logical functional division, and another division method can be used in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0107] In addition, each functional unit / module in each embodiment of the application can be integrated in one unit / module, or each unit / module can exist physically, or two or more units / modules can be integrated together. The integrated unit / module can be realized in the form of hardware or in the form of a software program module.
[0108] If the integrated units / modules are implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the storage unit can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0109] If the integrated units / modules are implemented in the form of software program modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the essential part or all or part of the technical solutions that make contributions to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0110] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0111] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.
[0112] It is to be understood that the application is not limited to the precise construction herein disclosed and shown in the drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application is limited only by the claims that follow.
Claims
1. A method of named entity recognition, characterized by, The method comprises the following steps: obtaining multi-modal financial data to be identified, wherein the multi-modal financial data comprises financial text and financial images associated with the financial text; obtaining a financial image feature vector based on the financial images; identifying a financial noun in the financial images by using a large language model based on the financial image feature vector and financial knowledge; converting words in the financial text into word vectors; splicing the financial noun and the word vectors to obtain a vector sequence; identifying a named entity in the financial text based on the vector sequence by using a named entity model.
2. The method of claim 1, wherein, The method comprises the following steps: performing image segmentation on the financial images to obtain at least one sub-image containing financial key data; extracting a financial feature vector from the sub-image by using an image feature extraction model, wherein the financial feature vector comprises visual features of financial elements; performing normalization processing on the financial feature vector to obtain the financial image feature vector.
3. The method of claim 2, wherein, The method comprises the following steps: converting the financial images into grayscale images; performing noise reduction processing on the grayscale images; performing image enhancement processing on the grayscale images after noise reduction processing; performing image segmentation on the grayscale images after image enhancement processing to obtain at least one sub-image containing financial key data.
4. The method of claim 1, wherein, The method comprises the following steps: respectively encoding the financial image feature vector and the financial knowledge; splicing the encoded financial image feature vector and the encoded financial knowledge to obtain a fusion feature vector; inputting the fused feature vector and a prompt text related to a business type corresponding to the financial images into a large model to obtain a financial noun in the financial images.
5. The method according to any one of claims 1 to 4, characterized in that, After identifying the financial noun in the financial images, the method further comprises the following steps: performing format checking on the financial noun by using a format rule; performing validity checking on the financial noun that passes the format checking; retaining the financial noun that passes the validity checking; The method comprises the following steps: splicing the checked financial noun and the word vectors to obtain the vector sequence.
6. The method according to any one of claims 1 to 4, characterized in that, After identifying the named entity in the financial text based on the vector sequence by using the named entity model, the method further comprises the following steps: if there is a first named entity with multiple labels in the identified named entity, removing redundant labels in the first named entity; and / or if there is a second named entity with a label that does not meet the named entity rules in the financial field in the identified named entity, correcting the label of the second named entity; and / or if there is a third named entity with an error label in the identified named entity, correcting the label of the third named entity.
7. A named entity recognition apparatus characterized by comprising: The device comprises: an obtaining module, configured to obtain multi-modal financial data to be identified, wherein the multi-modal financial data comprises financial text and financial images associated with the financial text; The feature vector identification module is configured to acquire a financial image feature vector based on the financial image. The financial noun identification module is configured to identify a financial noun in the financial image based on the financial image feature vector and financial knowledge by using a large language model. The word vector conversion module is configured to convert a word in the financial text into a word vector. The splicing module is configured to splice the financial noun and the word vector to obtain a vector sequence. The named entity recognition module is configured to identify a named entity in the financial text based on the vector sequence by using a named entity model.
8. An electronic device, comprising: The method comprises: a processor and a memory connected to the processor in communication; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method of any one of claims 1 to 6.
10. A computer program product, characterised in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 6.