Data reporting method, device, computer equipment and medium based on artificial intelligence
Through an artificial intelligence-based data reporting method, character recognition models are used to extract and integrate regulatory message information, which solves the problems of low data reporting efficiency and low accuracy in existing technologies and realizes efficient and accurate data reporting logic updates.
Patent Information
- Application Number
- CN202310835402.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-07-07
AI Technical Summary
When faced with diverse regulatory requirements, existing data reporting methods result in repeated parsing and generation of logic, which consumes human resources, is inefficient and inaccurate, and poses compliance risks.
An artificial intelligence-based data reporting method is adopted to extract regulatory message information through a trained character recognition model, determine the baseline table and merge it with the newly added reporting table to form an updated reporting table for data reporting, and reuse the baseline table logic for updating.
It improves the accuracy and efficiency of data reporting, enhances the scalability of logical construction and data uniformity, reduces workload, and meets new regulatory requirements.
Smart Images

Figure CN117033371B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based data reporting method, device, computer equipment and medium. Background Art
[0002] At present, with the gradual increase in the intensity of financial supervision, the regulatory structure's requirements for data reporting are becoming more and more frequent. The existing method usually parses the messages according to different data reporting requirements, then generates reporting logic based on the parsing results, and processes data reporting according to the reporting logic to meet diverse data reporting requirements.
[0003] However, when faced with a large number of data reporting requirements, repeated message parsing and reporting logic generation will consume a lot of human resources and make the data reporting efficiency low. Moreover, data reporting under different reporting logics will lead to the inability to adapt to each other between data reporting results. There may be compliance risks during cross-checking, resulting in low accuracy of data reporting. Therefore, how to improve the accuracy and efficiency of data reporting has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a data reporting method, apparatus, computer equipment, and medium based on artificial intelligence to solve the problem of low accuracy and efficiency in data reporting.
[0005] In a first aspect, an embodiment of the present invention provides a data reporting method based on artificial intelligence, the data reporting method comprising:
[0006] Obtaining a reference regulatory message, inputting the reference regulatory message into a trained character recognition model to extract text information, and determining the extraction result as a reference reporting form;
[0007] Obtaining at least one stored current submission form, matching the reference submission form with all current submission forms respectively, obtaining a matching degree value between the reference submission form and each current submission form respectively, and determining the current submission form corresponding to the maximum value among all matching degree values as a baseline form, wherein the baseline form includes at least one first submission item;
[0008] Obtaining a newly added regulatory message, inputting the newly added regulatory message into the trained character recognition model to extract text information, and determining that the extraction result is a newly added reporting form, wherein the newly added reporting form includes at least one second reporting item;
[0009] Merging all first reporting items in the baseline table and all second reporting items in the newly added reporting table to obtain N target reporting items, wherein any two of the N target reporting items represent different reporting contents, and N is an integer greater than zero;
[0010] An updated reporting table is formed according to the N target reporting items, and the updated reporting table is used to report the acquired target data.
[0011] In a second aspect, an embodiment of the present invention provides an artificial intelligence-based data reporting device, the data reporting device comprising:
[0012] A first message extraction module is used to obtain a reference regulatory message, input the reference regulatory message into a trained character recognition model to extract text information, and determine the extraction result as a reference reporting table;
[0013] a baseline table determination module, configured to obtain at least one stored current submission table, match the reference submission table with all current submission tables respectively, obtain a matching degree value between the reference submission table and each current submission table respectively, and determine the current submission table corresponding to the maximum matching degree value among all the matching degree values as the baseline table, wherein the baseline table includes at least one first submission item;
[0014] a second message extraction module, configured to obtain a newly added regulatory message, input the newly added regulatory message into the trained character recognition model to extract text information, and determine that the extraction result is a newly added reporting form, wherein the newly added reporting form includes at least one second reporting item;
[0015] a reporting item fusion module, configured to fuse all first reporting items in the baseline table and all second reporting items in the newly added reporting table to obtain N target reporting items, wherein any two of the N target reporting items represent different reporting contents, and N is an integer greater than zero;
[0016] The data reporting module is used to form an updated reporting table according to the N target reporting items, and use the updated reporting table to report the acquired target data.
[0017] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data reporting method as described in the first aspect when executing the computer program.
[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data reporting method as described in the first aspect is implemented.
[0019] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0020] Obtain a reference regulatory message, input the reference regulatory message into the trained character recognition model to extract text information, determine the extraction result as a reference reporting table, obtain at least one stored current reporting table, match the reference reporting table with all current reporting tables respectively, obtain the matching degree value between the reference reporting table and each current reporting table respectively, determine the current reporting table corresponding to the maximum value of all matching degree values as the baseline table, the baseline table includes at least one first reporting item, obtain a new regulatory message, input the new regulatory message into the trained character recognition model to extract text information, determine the extraction result as a new reporting table, the new reporting table includes at least one The second reporting item merges all the first reporting items in the baseline table and all the second reporting items in the newly added reporting table to obtain N target reporting items. The reporting contents represented by any two target reporting items among the N target reporting items are different. An updated reporting table is formed according to the N target reporting items, and the updated reporting table is used to report the acquired target data to determine the baseline table for data reporting. When facing new regulatory messages, the reporting logic of the baseline table can be reused. It is only necessary to update the baseline table to obtain an updated reporting table that meets the new regulatory messages, which improves the scalability of the reporting logic construction and the data uniformity, thereby improving the accuracy and efficiency of data reporting. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 This is a schematic diagram of an application environment of an artificial intelligence-based data reporting method provided in the first embodiment of the present invention;
[0023] Figure 2 This is a flow chart of a data reporting method based on artificial intelligence provided in the first embodiment of the present invention;
[0024] Figure 3 This is a structural diagram of an artificial intelligence-based data reporting device provided in the second embodiment of the present invention;
[0025] Figure 4 This is a structural diagram of a computer device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0026] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.
[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0028] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0030] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0032] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0033] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0034] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0035] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0036] The first embodiment of the present invention provides an artificial intelligence-based data reporting method that can be applied in Figure 1 The application environment in which the client communicates with the server. The client includes but is not limited to computer devices such as PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0037] See also Figure 2 , is a flow chart of a data reporting method based on artificial intelligence provided by the first embodiment of the present invention. The above data reporting method can be applied to Figure 1The client in the client, the computer device corresponding to the client connects to the server to obtain the reference supervision message, the new supervision message and at least one current reporting table stored. The computer device corresponding to the client can receive the user's data reporting instructions and the target data that needs to be reported. The computer device corresponding to the client is deployed with a trained character recognition model, which can be used to extract characters from the reference supervision message and the new supervision message. Figure 2 As shown, the data reporting method may include the following steps:
[0038] Step S201: obtain a reference supervision message, input the reference supervision message into a trained character recognition model to extract text information, and determine the extraction result as a reference reporting table.
[0039] Among them, the reference regulatory message may refer to the regulatory message obtained within a historical time period, the regulatory message may refer to the message containing the regulatory agency's requirements for data reporting, the trained character recognition model may refer to the optical character recognition model, the trained character recognition model can be used to extract character information from an image containing characters, and the reference reporting form may refer to a form containing the regulatory reporting requirements in the reference regulatory message.
[0040] Specifically, the historical time period may refer to a historical document issuance time period of the regulatory agency. In this embodiment, the reference regulatory message may be represented in an image format, and the reference regulatory message in image format may be obtained through image scanning.
[0041] The trained character recognition model may include a trained character positioning sub-model and a trained character classification sub-model. The trained character positioning sub-model may include a trained first encoder and a trained first fully connected layer. The trained character positioning sub-model may be used to locate characters in the image corresponding to the reference regulatory message. The input of the trained character positioning sub-model may be the image corresponding to the reference regulatory message. The output of the trained character positioning sub-model may refer to the bounding box of the character in the image corresponding to the reference regulatory message. The bounding box may be represented by the coordinate points of the upper left corner and the lower right corner of the bounding box. According to the bounding box, the position information of the character in the image corresponding to the reference regulatory message may be obtained, thereby realizing the positioning of the character.
[0042] The trained character classification sub-model may include a trained second encoder and a trained second fully connected layer. The trained character classification sub-model may be used to classify characters within a bounding box. The input of the trained character classification sub-model may be a character capture sub-image, which may be obtained by cropping the image corresponding to the reference regulatory message through the bounding box. It should be noted that after obtaining the cropped image, a size normalization operation needs to be performed. In this embodiment, the image corresponding to the reference regulatory message may be directly used as the normalized size, and the pixel values of the pixels outside the bounding box may be set to zero to achieve a size normalization operation for inputting the trained character classification sub-model.
[0043] The output of the trained character classification sub-model can refer to the classification prediction value vector of the characters in the character interception sub-image. The classification probability vector can include several preset characters and their corresponding character classification prediction values. The classification prediction value vector is processed by a normalized exponential function, and the maximum value of the processing results corresponding to each character classification prediction value is taken. The preset character corresponding to the maximum value is used as the classification result of the character in the character interception sub-image.
[0044] After the reference regulatory message is input into the trained character recognition model for text information extraction, the extraction results of each character sub-image in the image corresponding to the reference regulatory message can be template matched with the preset character template to determine the semantic information corresponding to the extraction results of each character sub-image, such as the preset character template, such as table name, table Chinese name, table number, remarks, data item name, data item code, data item type, data element code, data item description, etc.
[0045] Template matching can be performed using the Euclidean distance method, that is, for the extraction result of any character interception sub-image, the Euclidean distance calculation is performed on the extraction result and each preset character template respectively to obtain the Euclidean distance value between each preset character template and the extraction result, and the minimum value of all Euclidean distance values is determined. If the minimum value is less than the preset distance threshold, the preset character template corresponding to the minimum value is determined to be the semantic information of the extraction result. In this embodiment, the preset distance threshold can be set to 10.
[0046] After template matching, the extraction results of the successfully matched character interception sub-image are determined as attribute characters, and the extraction results that do not belong to attribute characters in the image corresponding to the reference regulatory message are determined as content characters. For any attribute character, in the image corresponding to the reference regulatory message, the closest content characters are searched for the enclosing box corresponding to the attribute character. For example, the content characters searched for the "table name" attribute character are "JGGQXXB", the content characters searched for the "Chinese name in the table" attribute character are "institutional equity information table", the content characters searched for the "data item" attribute character are "serial number", and the content characters searched for the "data item code" attribute character are "LSH", etc. It is determined that the searched content characters and the attribute characters constitute a single table item, and all attribute characters are traversed to obtain all table items, and all the obtained table items constitute a reference reporting table.
[0047] It should be noted that the reference reporting table may include the "null rate" attribute character and the "NULL rate" attribute character. The content character corresponding to the "null rate" attribute character is zero when the reporting requirement cannot be null and the actual data is not null. When the reporting requirement can be null, the non-null percentage is calculated based on the actual data, and the content character is a positive number at this time. When the reporting requirement cannot be null, the non-null percentage is calculated based on the actual data, and the content character is a negative number at this time.
[0048] Similarly, the content character corresponding to the above-mentioned "NULL rate" attribute character is zero when the reporting requirement cannot be NULL and the actual data is not NULL. When the reporting requirement can be NULL, the non-NULL percentage is calculated based on the actual data, and the content character is a positive number at this time. When the reporting requirement cannot be NULL, the non-NULL percentage is calculated based on the actual data, and the content character is a negative number at this time.
[0049] Optionally, the reference regulatory message is input into a trained character recognition model to extract text information, and the extraction result is determined as a reference submission table including:
[0050] Inputting the reference supervision message into the trained character recognition model to extract text information and obtain a first text extraction result;
[0051] The first positioning identifier and the first frequency of the reference regulatory message are input, and a reference reporting table is formed according to the first character extraction result, the first positioning identifier and the first frequency.
[0052] Among them, the first text extraction result may refer to all character information extracted from the reference regulatory message, the first positioning identifier may refer to the positioning identifier of the storage location corresponding to the reference regulatory message, and the first frequency may refer to the data reporting frequency indicated by the reference regulatory message. The data reporting frequency may include hours, days, weeks, months, years, etc.
[0053] Specifically, with "positioning identifier" as the attribute character and the first positioning identifier as the content character corresponding to the "positioning identifier" attribute character, the positioning identifier table item can be obtained from the attribute character and the content character. Similarly, with "frequency" as the attribute character and the first frequency as the content character corresponding to the "frequency" attribute character, the frequency identifier table item can be obtained from the attribute character and the content character. The positioning identifier table item, the frequency identifier table item and all the table items corresponding to the first text extraction result can be spliced together to obtain a reference submission table in accordance with the preset attribute character order. For example, for the three attribute characters of "positioning identifier", "frequency" and "data item", the preset attribute character order can be "data item" in the first place, "positioning identifier" in the second place and "frequency" in the third place.
[0054] In this embodiment, a reference reporting table is formed by the first text extraction result, the first positioning identifier and the first frequency, which ensures the uniformity of the reporting table format, facilitates comparison between subsequent reporting tables, and at the same time increases the richness of the content of the reporting table, improves the comparison accuracy when comparing subsequent reporting tables, thereby improving the efficiency and accuracy of data reporting.
[0055] The above steps of obtaining the reference regulatory message, inputting the reference regulatory message into the trained character recognition model to extract text information, and determining the extraction result as the reference reporting table, extracting the text information of the reference regulatory message, and providing basic text information for the subsequent determination of the baseline table, so as to determine the appropriate baseline table and improve the accuracy of data reporting.
[0056] Step S202: Obtain at least one stored current reporting table, match the reference reporting table with all current reporting tables respectively, obtain the matching degree values between the reference reporting table and each current reporting table respectively, and determine the current reporting table corresponding to the maximum value among all matching degree values as the baseline table.
[0057] Among them, the current reporting form may refer to the reporting form that is currently used for data reporting, the matching degree value may be used to characterize the similarity between the reference reporting form and the corresponding current reporting form, and the baseline table may refer to the reporting form that provides benchmark information for the final reporting form. The baseline table may include at least one first reporting item, and the first reporting item is also the above-mentioned table item.
[0058] Specifically, in this embodiment, after obtaining the current reporting table corresponding to the maximum value among all matching degree values, the current reporting table is directly used as the baseline table.
[0059] In one embodiment, the union of the reference submission table and the current submission table corresponding to the maximum value of all matching degree values can be taken as the baseline table. At this time, the content characters in the table items in the baseline table that belong to the intersection of the two table items shall be based on the reference submission table.
[0060] Optionally, matching the reference submission form with all current submission forms respectively to obtain a matching degree value between the reference submission form and each current submission form, and determining the current submission form corresponding to the maximum matching degree value among all the matching degree values as the baseline form includes:
[0061] For any current submission form, use the Euclidean distance to match the reference submission form with the current submission form to obtain the matching degree value corresponding to the current submission form;
[0062] All current submission tables are traversed to obtain the matching degree values corresponding to the current submission tables, and the current submission table corresponding to the maximum value of all matching degree values is determined as the baseline table.
[0063] Among them, the Euclidean distance can be used to calculate the difference between two vectors, and the reference submission form and the current submission form can both be regarded as 2*N-dimensional vectors, that is, vectors with two rows and N columns.
[0064] Specifically, when calculating the Euclidean distance, in this embodiment, it is assumed that the attribute characters in each table item of different reporting tables are consistent, so the Euclidean distance calculation can be performed only on the content characters. It should be noted that the implementer can set a pre-comparison strategy to compare the attribute characters of each table item in different reporting tables. If there are inconsistent attribute characters, the table item corresponding to the attribute character can be ignored.
[0065] When calculating the Euclidean distance, the content characters need to be encoded. The encoding process uses a preset encoding mapping table for encoding. The encoding mapping table can include a mapping relationship between content characters and encoding values. The encoding values are all in digital form, for example, hexadecimal values can be used to ensure that the Euclidean distance can be calculated quickly.
[0066] The difference between the code values of the content characters corresponding to the positions in the reference submission table and the current submission table is calculated to obtain the difference corresponding to each position, and the square root of the sum of all the differences is taken as the calculation result of the Euclidean distance.
[0067] It should be noted that different weights can be set for different positions. When calculating the sum of squares of all differences, the squares of the corresponding differences at different positions are multiplied by the corresponding weights and then the sum is calculated. The weight can be determined according to the position, that is, the attribute character corresponding to the position. For example, if the attribute character is "data item", the weight can be set to 0.6, if the attribute character is "frequency", the weight can be set to 0.1, if the attribute character is "data item type", the weight can be set to 0.1, if the attribute character is "empty rate", the weight can be set to 0.1, if the attribute character is "NULL rate", the weight can be set to 0.1, and so on.
[0068] In this embodiment, the difference between the reference reporting form and the current reporting form is measured by Euclidean distance, so as to determine the current reporting form that is closest to the reference reporting form, and then determine the baseline table, providing a baseline table for the update of subsequent data reporting tables, ensuring the update efficiency of the baseline table, and thus improving the accuracy and efficiency of data reporting.
[0069] The above-mentioned step of obtaining at least one stored current reporting form, matching the reference reporting form with all the current reporting forms respectively, obtaining the matching degree values between the reference reporting form and each current reporting form respectively, and determining the current reporting form corresponding to the maximum value of all matching degree values as the baseline form, can determine the baseline form based on the obtained reference reporting form, so that the determination of the baseline form is applicable to the data relationship between different regulatory agencies and different regulatory messages, thereby improving the accuracy and efficiency of subsequent data reporting.
[0070] Step S203: Acquire a newly added supervision message, input the newly added supervision message into a trained character recognition model to extract text information, and determine that the extraction result is a newly added reporting form.
[0071] The newly added supervision message may refer to the most recently acquired supervision message, the newly added reporting table includes at least one second reporting item, and the second reporting item may refer to an item in the newly added reporting table.
[0072] Specifically, after the newly added regulatory message is input into the trained character recognition model for text information extraction, the extraction results of each character sub-image in the image corresponding to the newly added regulatory message can be template matched with the preset character template to determine the semantic information corresponding to the extraction results of each character sub-image.
[0073] After template matching, the extraction results of the successfully matched character interception sub-image are determined as attribute characters, and the extraction results that do not belong to attribute characters in the image corresponding to the newly added regulatory message are determined to be content characters. For any attribute character, in the image corresponding to the newly added regulatory message, the closest content character is searched for the enclosing box corresponding to the attribute character, and it is determined that the searched content character and the attribute character form a single table item. All attribute characters are traversed to obtain all table items, and all the obtained table items are used to form a new reporting table. That is, the construction method of the new reporting table is consistent with the reference reporting table.
[0074] Optionally, the newly added regulatory message is input into a trained character recognition model to extract text information, and the extraction result is determined to be a newly submitted form including:
[0075] Input the newly added supervision message into the trained character recognition model to extract text information and obtain a second text extraction result;
[0076] Enter the second positioning identifier and second frequency of the newly added regulatory message, and form a new reporting table based on the second text extraction result, the second positioning identifier and the second frequency.
[0077] Among them, the second text extraction result may refer to all character information extracted from the newly added regulatory message, the second location identifier may refer to the location identifier of the storage location corresponding to the newly added regulatory message, and the second frequency may refer to the data reporting frequency indicated by the reference regulatory message.
[0078] Specifically, corresponding positioning identifier table entries and frequency identifier table entries are also constructed based on the positioning identifier and frequency. The positioning identifier table entries, frequency identifier table entries and all table entries corresponding to the first text extraction result can be spliced together to obtain a new reporting table according to the preset attribute character order.
[0079] In this embodiment, a new reporting form is formed by the second text extraction result, the second positioning identifier and the second frequency, which ensures the uniformity of the reporting form format, facilitates comparison between subsequent reporting forms, and at the same time increases the richness of the content of the reporting form, improves the comparison accuracy when comparing subsequent reporting forms, thereby improving the efficiency and accuracy of data reporting.
[0080] The above-mentioned steps of obtaining the newly added regulatory message, inputting the newly added regulatory message into the trained character recognition model to extract text information, and determining the extraction result as the newly added reporting form, extracting the text information of the newly added regulatory message, and providing the information of the newly added regulatory message for the subsequent update of the reporting form, thereby improving the accuracy of data reporting.
[0081] Step S204: merge all first submission items in the baseline table and all second submission items in the newly added submission table to obtain N target submission items.
[0082] Among them, the target reporting item may refer to the reporting table item obtained after fusion, N is an integer greater than zero, and the reporting contents represented by any two target reporting items among the N target reporting items are different.
[0083] Optionally, all first reporting items in the baseline table and all second reporting items in the newly added reporting table are merged to obtain N target reporting items including:
[0084] For any first submission item, match the first submission item with each second submission item respectively;
[0085] If there is a second submission item that matches the first submission item, the first submission item is determined to be the target submission item, and each first submission item is traversed to obtain M target submission items;
[0086] Determine N target reporting items based on the M target reporting items and all second reporting items in the newly added reporting table.
[0087] The matching result may represent the degree of similarity between the corresponding first reported item and the corresponding second reported item. The matching result may include consistency and inconsistency, and M is an integer greater than zero.
[0088] Specifically, the Euclidean distance calculation method can still be used for matching. Similarly, the reported items also need to be encoded.
[0089] If there is a second submission item and the matching result of the first submission item is consistent, it means that the second submission item and the first submission item are repeated. Only one of the submission items needs to be retained as the target submission item, that is, the first submission item is retained as the target submission item.
[0090] In this embodiment, each first reported item is matched with each second reported item respectively, and each first reported item is determined as a target reported item based on the matching result, thereby avoiding duplication of reported items and improving the efficiency of subsequent data reporting.
[0091] Optionally, after matching the first submitted item with each second submitted item respectively, the method further includes:
[0092] According to the preset scoring rules, the scoring results of the first submitted item and each second submitted item are obtained, and each scoring result is compared with the preset scoring threshold to obtain a comparison result of the corresponding scoring results;
[0093] If any comparison result is that the corresponding score result is greater than the score threshold, it is determined that there is a matching result of the second submission item and the first submission item that is consistent.
[0094] Among them, the preset scoring rules may include a mapping relationship between the matched Euclidean distance and the scoring result, the scoring result can be used to characterize the degree of similarity between the corresponding first report item and the corresponding second report item, and the scoring threshold can be used to measure whether the corresponding first report item and the corresponding second report item are sufficiently similar.
[0095] Specifically, in this embodiment, the value range of the scoring result is [0,1], and the preset scoring threshold can be set to 0.95, so as to ensure that reporting items with the same semantics but different characters are also considered similar, such as "order serial number" and "order serial number". It should be noted that the implementer can adjust the scoring threshold according to actual conditions.
[0096] In one embodiment, different scoring thresholds may be set according to the attribute characters of different reported items, thereby increasing the flexibility of the comparison process.
[0097] In this embodiment, the comparison result is determined by mapping the scores using Euclidean distance and comparing them with a preset score threshold, thereby improving the flexibility of the comparison process, avoiding misjudgment, and improving the accuracy of subsequent data reporting.
[0098] Optionally, based on the M target submission items and all second submission items in the newly added submission table, the N target submission items are determined to include:
[0099] Among all the second submission items in the newly added submission table, the second submission item that matches any target submission item is determined as the reference submission item;
[0100] Determine all second submission items except the reference submission items in the newly added submission table as target submission items, and obtain K target submission items;
[0101] N target reporting items are composed of M target reporting items and K target reporting items.
[0102] Among them, the reference submission item may refer to a second submission item that matches the first submission item, that is, a second submission item that does not need to be a target submission item, and K is an integer greater than zero.
[0103] Specifically, all second submission items in the newly added submission table except the reference submission items are new information. Therefore, all second submission items in the newly added submission table except the reference submission items are used as target submission items. M target submission items and K target submission items constitute N target submission items. Obviously, N satisfies N=M+K.
[0104] In this embodiment, the second reporting item that has not been successfully matched is determined as the target reporting item to ensure that no new reporting requirements are missed, thereby improving the accuracy of subsequent data reporting.
[0105] The above steps of fusing all the first reporting items in the baseline table and all the second reporting items in the newly added reporting table to obtain N target reporting items can maximize the use of existing reporting resources when processing the newly added regulatory reporting table, quickly determine the target reporting items, greatly improve the scalability of the reported data, and thus improve the efficiency of data reporting.
[0106] Step S205 , forming an updated reporting table according to the N target reporting items, and using the updated reporting table to report the acquired target data.
[0107] The updated reporting table may refer to a reporting table used for reporting target data, and the target data may refer to data that currently needs to be reported.
[0108] Specifically, the N target reporting items are spliced according to their corresponding attribute characters and in a preset attribute character sequence to obtain an updated reporting table.
[0109] The above steps of forming an updated reporting table based on N target reporting items and using the updated reporting table to report the acquired target data maximize the reuse of the existing reporting logic. After multiple data reporting processes, the baseline table can basically cover the reported data, greatly improving the response time for data reporting of target data, reducing the workload of data reporting, and being able to comply with new regulatory requirements, thereby improving the accuracy and efficiency of data reporting.
[0110] In this embodiment, a baseline table for data reporting is determined. When faced with new regulatory messages, the reporting logic of the baseline table can be reused. Only the baseline table needs to be updated to obtain an updated reporting table that meets the new regulatory messages. This improves the scalability and data uniformity of the reporting logic construction, thereby improving the accuracy and efficiency of data reporting.
[0111] Corresponding to the data reporting method based on artificial intelligence in the above embodiment, Figure 3 The structural block diagram of the artificial intelligence-based data reporting device provided in the second embodiment of the present invention is shown. The above-mentioned data reporting device is applied to the client, and the computer device corresponding to the client is connected to the server to obtain reference regulatory messages, new regulatory messages and at least one stored current reporting table. The computer device corresponding to the client can receive the user's data reporting instructions and the target data that needs to be reported. The computer device corresponding to the client is deployed with a trained character recognition model, which can be used to extract characters from the reference regulatory messages and the new regulatory messages. For ease of explanation, only the parts related to the embodiment of the present invention are shown.
[0112] See also Figure 3 , the data reporting device includes:
[0113] A first message extraction module 31 is used to obtain a reference regulatory message, input the reference regulatory message into a trained character recognition model to extract text information, and determine the extraction result as a reference reporting table;
[0114] A baseline table determination module 32 is configured to obtain at least one stored current submission table, match the reference submission table with all current submission tables, obtain a matching degree value between the reference submission table and each current submission table, and determine the current submission table corresponding to the maximum matching degree value among all the matching degree values as the baseline table, wherein the baseline table includes at least one first submission item;
[0115] A second message extraction module 33 is configured to obtain a newly added regulatory message, input the newly added regulatory message into a trained character recognition model to extract text information, and determine that the extraction result is a newly added reporting form, where the newly added reporting form includes at least one second reporting item;
[0116] A report item fusion module 34 is configured to fuse all first report items in the baseline table and all second report items in the newly added report table to obtain N target report items, wherein any two of the N target report items represent different report contents, and N is an integer greater than zero;
[0117] The data reporting module 35 is used to form an updated reporting table according to the N target reporting items, and use the updated reporting table to report the acquired target data.
[0118] Optionally, the first message extraction module 31 includes:
[0119] A first character extraction unit is configured to input the reference supervision message into a trained character recognition model to extract character information and obtain a first character extraction result;
[0120] The first reporting table forming unit is used to input the first positioning identifier and the first frequency of the reference supervision message, and form a reference reporting table according to the first character extraction result, the first positioning identifier and the first frequency.
[0121] Optionally, the baseline table determination module 32 includes:
[0122] A report form matching unit is used to match the reference report form with the current report form using the Euclidean distance for any current report form, and obtain a matching degree value corresponding to the current report form;
[0123] The reporting table traversal unit is used to traverse all current reporting tables, obtain the matching degree values corresponding to the current reporting tables, and determine that the current reporting table corresponding to the maximum value of all matching degree values is the baseline table.
[0124] Optionally, the second message extraction module 33 includes:
[0125] A second character extraction unit is used to input the newly added supervision message into the trained character recognition model to extract character information and obtain a second character extraction result;
[0126] The second reporting table forming unit is used to input the second positioning identifier and the second frequency of the newly added regulatory message, and form a newly added reporting table according to the second text extraction result, the second positioning identifier and the second frequency.
[0127] Optionally, the reporting item fusion module 34 includes:
[0128] a submission item matching unit, configured to match the first submission item with each second submission item for each first submission item;
[0129] A submission item traversal unit is configured to determine that the first submission item is a target submission item if there is a second submission item that matches the first submission item, and traverse each first submission item to obtain M target submission items, where M is an integer greater than zero;
[0130] The reporting item determination unit is used to determine N target reporting items based on the M target reporting items and all second reporting items in the newly added reporting table.
[0131] Optionally, the reporting item fusion module 34 further includes:
[0132] a matching scoring unit, configured to obtain a scoring result of each of the first submitted item and each of the second submitted items according to a preset scoring rule, and compare each scoring result with a preset scoring threshold to obtain a comparison result of the corresponding scoring results;
[0133] The matching result determining unit is configured to determine that there is a matching result of the second submission item that is consistent with the first submission item if any comparison result indicates that the corresponding scoring result is greater than a scoring threshold.
[0134] Optionally, the reporting item determination unit includes:
[0135] A submission item screening subunit is used to determine, among all the second submission items in the newly added submission table, a second submission item that matches any target submission item as a reference submission item;
[0136] a submission item determination subunit, configured to determine all second submission items in the newly added submission table except the reference submission item as target submission items, to obtain K target submission items, where K is an integer greater than zero;
[0137] The submission item combination subunit is used to form N target submission items from M target submission items and K target submission items.
[0138] It should be noted that the information interaction, execution process, etc. between the above-mentioned modules, units, and sub-units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0139] Figure 4 This is a schematic diagram of the structure of a computer device provided in the third embodiment of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned data reporting method embodiments are implemented.
[0140] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 4 The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.
[0141] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.
[0142] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.
[0143] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0144] The present invention may implement all or part of the processes in the above-mentioned method embodiments, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0145] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0146] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0147] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0148] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0149] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A data reporting method based on artificial intelligence, characterized in that: The data reporting method includes: Obtaining a reference regulatory message, inputting the reference regulatory message into a trained character recognition model to extract text information, and determining the extraction result as a reference reporting form; Obtaining at least one stored current submission form, matching the reference submission form with all current submission forms respectively, obtaining a matching degree value between the reference submission form and each current submission form respectively, and determining the current submission form corresponding to the maximum value among all matching degree values as a baseline form, wherein the baseline form includes at least one first submission item; Obtaining a newly added regulatory message, inputting the newly added regulatory message into the trained character recognition model to extract text information, and determining that the extraction result is a newly added reporting form, wherein the newly added reporting form includes at least one second reporting item; Merging all first reporting items in the baseline table and all second reporting items in the newly added reporting table to obtain N target reporting items, wherein any two of the N target reporting items represent different reporting contents, and N is an integer greater than zero; An updated reporting table is formed according to the N target reporting items, and the updated reporting table is used to report the acquired target data.
2. The data reporting method according to claim 1, characterized in that: The step of inputting the reference supervision message into a trained character recognition model to extract text information, and determining the extraction result as a reference reporting table comprises: Inputting the reference supervision message into a trained character recognition model to extract text information, thereby obtaining a first text extraction result; The first location identifier and the first frequency of the reference regulatory message are input, and the reference reporting table is formed according to the first character extraction result, the first location identifier and the first frequency.
3. The data reporting method according to claim 1, wherein: Matching the reference submission table with all current submission tables respectively to obtain a matching degree value between the reference submission table and each current submission table, and determining the current submission table corresponding to the maximum value among all matching degree values as the baseline table includes: For any current reporting form, the reference reporting form and the current reporting form are matched using the Euclidean distance to obtain a matching degree value corresponding to the current reporting form; All current reporting tables are traversed to obtain matching degree values corresponding to the current reporting tables, and the current reporting table corresponding to the maximum value among all matching degree values is determined to be the baseline table.
4. The data reporting method according to claim 1, wherein: The step of inputting the newly added regulatory message into the trained character recognition model to extract text information, and determining the extraction result as a newly added reporting form comprises: Inputting the newly added supervision message into the trained character recognition model to extract text information, thereby obtaining a second text extraction result; The second location identifier and the second frequency of the newly added regulatory message are entered, and the newly added reporting table is formed according to the second character extraction result, the second location identifier and the second frequency.
5. The data reporting method according to any one of claims 1 to 4, characterized in that: The fusion of all first reporting items in the baseline table and all second reporting items in the newly added reporting table to obtain N target reporting items includes: For any first submission item, match the first submission item with each second submission item respectively; If there is a second submission item that matches the first submission item, the first submission item is determined to be the target submission item, and each first submission item is traversed to obtain M target submission items, where M is an integer greater than zero; The N target reporting items are determined according to the M target reporting items and all second reporting items in the newly added reporting table.
6. The data reporting method according to claim 5, characterized in that: After matching the first submitted item with each second submitted item respectively, the method further includes: According to a preset scoring rule, obtaining scoring results for each of the first submitted item and each of the second submitted items, and comparing each scoring result with a preset scoring threshold to obtain a comparison result of the corresponding scoring results; If any comparison result is that the corresponding score result is greater than the score threshold, it is determined that there is a second submission item whose matching result is consistent with the first submission item.
7. The data reporting method according to claim 5, characterized in that: The determining of the N target reporting items according to the M target reporting items and all second reporting items in the newly added reporting table includes: Determine, among all the second submission items in the newly added submission table, the second submission item that matches any target submission item as a reference submission item; Determine all second submission items in the newly added submission table except the reference submission item as the target submission items, and obtain K target submission items, where K is an integer greater than zero; The M target reporting items and the K target reporting items constitute the N target reporting items.
8. A data reporting device based on artificial intelligence, characterized in that: The data reporting device includes: A first message extraction module is used to obtain a reference regulatory message, input the reference regulatory message into a trained character recognition model to extract text information, and determine the extraction result as a reference reporting table; a baseline table determination module, configured to obtain at least one stored current submission table, match the reference submission table with all current submission tables respectively, obtain a matching degree value between the reference submission table and each current submission table respectively, and determine the current submission table corresponding to the maximum matching degree value among all the matching degree values as the baseline table, wherein the baseline table includes at least one first submission item; a second message extraction module, configured to obtain a newly added regulatory message, input the newly added regulatory message into the trained character recognition model to extract text information, and determine that the extraction result is a newly added reporting form, wherein the newly added reporting form includes at least one second reporting item; a reporting item fusion module, configured to fuse all first reporting items in the baseline table and all second reporting items in the newly added reporting table to obtain N target reporting items, wherein any two of the N target reporting items represent different reporting contents, and N is an integer greater than zero; The data reporting module is used to form an updated reporting table according to the N target reporting items, and use the updated reporting table to report the acquired target data.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the data reporting method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data reporting method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data processing method and device, storage medium and electronic equipment
CN114218220A
Data processing method and device, equipment and storage medium
CN115601122A