An information checking method based on big data

By identifying and verifying keywords in user-submitted inquiry information, distinguishing between fuzzy and standard keywords, calculating fuzzy probability coefficients, verifying the category represented, and optimizing feedback methods, the problem of poor feedback caused by unclear expression of message information on the bank's online processing platform has been solved, achieving more accurate and efficient feedback.

CN117851590BActive Publication Date: 2025-10-21HUNAN SANXIANG BANK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311686619.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-10-21
Estimated Expiration
2043-12-11

AI Technical Summary

Technical Problem

In the online banking platform's message service, users' messages are often unclear, resulting in poor feedback.

Method used

By acquiring user query information, identifying keywords and comparing them with sample tags, distinguishing between fuzzy keywords and standard keywords, calculating fuzzy probability coefficients, verifying the category represented by the query information, and selecting feedback methods based on the verification results, including prioritizing the push of related feedback information or conducting secondary verification.

Benefits of technology

It improves the accuracy of message information analysis, reduces the probability of invalid feedback, and ensures the effectiveness of feedback information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117851590B_ABST
    Figure CN117851590B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data verification classification, and more particularly to an information verification method based on big data, which obtains inquiry information published by a user terminal on an interactive platform, compares keywords in the inquiry information with sample labels, identifies ambiguous keywords and standard keywords in the inquiry information, classifies the ambiguous keywords according to corresponding sub-labels of the ambiguous keywords, determines fuzzy probability coefficients of the ambiguous keywords of each category, determines a maximum fuzzy probability coefficient as a probability consistency coefficient, verifies a representation direction category of the inquiry information according to the probability consistency coefficient, selects a feedback mode for the inquiry information according to a verification result, considers the influence of the clarity of the message information on the verification accuracy through the above process, and feeds back the inquiry information according to the verification result, so as to reduce the probability of invalid feedback through data verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data verification and classification, and in particular to an information verification method based on big data. Background Art

[0002] With the popularization of the Internet and the rapid development of information technology, big data technology has become an important support in the information age. Using big data technology to collect, organize and analyze massive amounts of data provides important support for providing scientific and effective information verification results.

[0003] Chinese Patent Publication No.: CN111445212A, discloses the following content: This invention discloses an enterprise talent information management system based on big data, including an internal talent management module and an external talent management module; the external talent management module includes an external database, a collection unit, a classification unit, a retrieval unit, a comparison unit, and a recruitment unit, and the collection unit is used to collect external talent information data from the external database. In the present invention, the information confidentiality unit is set to keep the internal talent information of the enterprise confidential, thereby improving the security of the information management system; the talent information update unit is set to automatically update the talent information data and improve the efficiency of updating the talent information data; the talent information tracking unit is set to track and match the internal talent information of the enterprise, and it is possible to verify whether the information provided by the talent is consistent with the facts; the login verification unit is set to verify the user, preventing outsiders from entering the talent information management system at will.

[0004] However, the prior art still has the following problems:

[0005] In the prior art, when analyzing the message information in the message service opened by the bank's online processing platform, the user who left the message did not clearly express the business to be handled, resulting in inaccurate analysis of the true purpose of the user's message information, resulting in poor feedback effect to the user. Summary of the Invention

[0006] To this end, the present invention provides an information verification method based on big data, which is used to overcome the problem that when analyzing the message information in the message service opened by the bank's online processing platform, the user terminal leaving the message does not clearly express the business to be handled, resulting in inaccurate analysis of the real purpose of the user terminal's message information, resulting in poor feedback effect to the user terminal.

[0007] To achieve the above objectives, the present invention provides an information verification method based on big data, which includes:

[0008] Step S1: Obtain a query message posted by a user on the interactive platform, compare the keywords in the query message with sample tags, and identify fuzzy keywords and standard keywords in the query message. The sample tags each include a main tag and several sub-tags, and both the sub-tags and the main tag are keywords.

[0009] Step S2: if the query information does not contain standard keywords, classify the fuzzy keywords according to the sub-tags corresponding to the fuzzy keywords, determine the fuzzy probability coefficients of the fuzzy keywords in each category, and determine the maximum fuzzy probability coefficient as the probability consistency coefficient;

[0010] Step S3, checking the representation direction category of the inquiry information according to the probability consistency coefficient, where the representation direction category includes fuzzy representation direction and clear representation direction;

[0011] Step S4, selecting a feedback method for the inquiry information according to the verification result, including:

[0012] Identify the sample characterization tag corresponding to the query information, and give priority to pushing feedback information associated with the sample characterization tag;

[0013] Alternatively, the query information is subjected to a secondary check, including identifying fuzzy keywords in each text sentence of the query information, determining the association representation coefficient of each main tag based on the comparison results of the remaining keywords other than the fuzzy keywords in each text sentence with each sample sentence in the sample database, and giving priority to pushing feedback information associated with the main tag corresponding to the largest association representation coefficient.

[0014] Furthermore, the step S1 also includes pre-building sample labels, and the building process includes:

[0015] Obtain several sample sentences containing a single keyword from the sample database in advance, extract several remaining keywords from the sample sentences, and calculate the occurrence probability of each remaining keyword one by one.

[0016] If the probability of occurrence of the remaining keywords is greater than the predetermined probability threshold, the single keyword is used as the main tag, the remaining keywords are used as sub-tags of the main tag, and the probability of occurrence of the remaining keywords is determined as the association probability between the main tag and the sub-tag. The probability of occurrence of the remaining keywords is the probability of the remaining keywords appearing in the sample sentence.

[0017] Furthermore, in step S1, the process of identifying the fuzzy keywords and standard keywords in the query information based on the comparison results of the keywords in the query information and the sample labels includes:

[0018] If the keyword in the query information is the same as the main tag of the sample tag, the keyword is identified as a standard keyword;

[0019] If the keyword in the query information is the same as the sub-tag of the sample tag, the keyword is identified as an ambiguous keyword.

[0020] Furthermore, in step S2, the process of classifying the fuzzy keywords according to the sub-tags corresponding to the fuzzy keywords includes:

[0021] Based on the sub-tags corresponding to each fuzzy keyword, determine whether the categories of each fuzzy keyword are the same.

[0022] If the subtags corresponding to the fuzzy keywords belong to the same main tag, it is determined that the categories of the fuzzy keywords are the same.

[0023] Furthermore, in step S2, the fuzzy probability coefficient of each category of fuzzy keywords is determined, wherein,

[0024] According to formula (1), the fuzzy probability coefficient Ki corresponding to the fuzzy keyword belonging to category i is calculated.

[0025] ,

[0026] In formula (1), N0 represents the total number of fuzzy keywords in the query information, n represents the number of fuzzy keywords belonging to the i-th category, Pa represents the association probability between the sub-tag corresponding to the a-th fuzzy keyword belonging to the i-th category and the associated main tag, and i and a are both integers greater than 0.

[0027] Furthermore, in step S3, the process of checking the representation pointing category of the query information according to the probability consistency coefficient includes:

[0028] Compare the probability consistency coefficient with a preset coefficient comparison threshold,

[0029] If the probability consistency coefficient is greater than or equal to the coefficient comparison threshold, verify and determine that the representation direction category of the inquiry information is clear representation direction;

[0030] If the probability consistency coefficient is less than the coefficient comparison threshold, it is checked and determined that the representation orientation category of the query information is fuzzy representation orientation.

[0031] Furthermore, the step S3 further includes determining the representation direction category of the query information based on the keywords in the query information,

[0032] If the query information contains standard keywords, it is determined that the representation orientation category of the query information is clear representation orientation.

[0033] Furthermore, in step S4, the process of selecting a feedback method for the inquiry information based on the verification result includes:

[0034] If the representation direction category of the query information is clear representation direction, identify the representation sample tag corresponding to the query information, and give priority to pushing feedback information associated with the main tag;

[0035] If the representation direction category of the query information is fuzzy representation direction, a secondary check is performed on the query information.

[0036] Furthermore, in step S4, the process of identifying the sample label corresponding to the query information includes:

[0037] If the query information contains a standard keyword, the main tag corresponding to the standard keyword is determined as the representative sample tag;

[0038] If there is no standard keyword in the query information, a target category fuzzy keyword is determined, and the main label of the sub-label corresponding to the target category fuzzy keyword is used as the representative sample label. The target category fuzzy keyword is a fuzzy keyword of the category corresponding to the probability consistency coefficient.

[0039] Furthermore, in step S4, the process of determining the association representation coefficient of each main tag according to the comparison results of the remaining keywords excluding the fuzzy keywords in each text sentence with each sample sentence in the sample database includes:

[0040] The occurrence probability of each of the remaining keywords in the sample sentence containing a single sub-tag is determined, and an average value of the occurrence probabilities is determined as the association representation coefficient of the main tag.

[0041] Compared with the prior art, the present invention obtains inquiry information posted by a user terminal on an interactive platform, compares the keywords in the inquiry information with sample labels, identifies fuzzy keywords and standard keywords in the inquiry information, classifies the fuzzy keywords according to the sub-labels corresponding to the fuzzy keywords, and determines the fuzzy probability coefficients of the fuzzy keywords in each category, determines the maximum fuzzy probability coefficient as the probability consistency coefficient, verifies the representation pointing category of the inquiry information according to the probability consistency coefficient, selects the feedback method for the inquiry information according to the verification result, considers the influence of the clarity of the message information on the verification accuracy through the above process, and provides feedback on the inquiry information according to the verification result, thereby reducing the probability of invalid feedback through data verification.

[0042] In particular, in the present invention, sample labels are pre-constructed. In actual situations, if a certain keyword appears in a sentence, the remaining keywords in the sentence have a high probability of appearing, indicating that this keyword and the remaining keywords are highly correlated in practice. Therefore, the keyword is used as the main label, and the remaining keywords with a high probability of appearing are used as sub-labels. By analyzing big data, the connection between the keywords can be reliably found, providing an important basis for subsequent verification of the inquiry information.

[0043] In particular, in the present invention, fuzzy keywords and standard keywords in the query information are identified based on the comparison results between the keywords in the query information and the sample labels. In actual situations, if the keywords in the query information are the same as the main labels, it indicates that the intention of the query information is clearly expressed. If the keywords in the query information are the same as the sub-labels, it indicates that the intention of the query information may not be clear and needs further verification. Therefore, the keywords that are the same as the main labels are determined as standard keywords, and the keywords that are the same as the sub-labels are determined as fuzzy keywords, so as to distinguish the keywords in the query information, further verify whether the intention expressed by the query information is clear, and feedback is given to the query information based on the verification results, thereby reducing the probability of invalid feedback through data verification.

[0044] In particular, in the present invention, fuzzy keywords are classified according to the sub-tags corresponding to the fuzzy keywords. In actual situations, if the main tags to which the sub-tags corresponding to the fuzzy keywords belong are the same, it indicates that the intentions expressed by the fuzzy keywords are the same. Therefore, the fuzzy keywords are regarded as the same category, which facilitates the subsequent verification of whether the intentions expressed in the query information are clear based on the situation of the fuzzy keywords in each category, so as to provide effective feedback on the query information through the verification results.

[0045] In particular, in the present invention, the fuzzy probability coefficient of each category of fuzzy keywords is determined. The fuzzy probability coefficient is calculated by the proportion of fuzzy keywords belonging to the same category in the fuzzy keywords in the query information and the average value of the association probability between the sub-tags corresponding to each fuzzy keyword belonging to this category and the associated main tag. In actual situations, the larger the proportion, the more similar the direction expressed by the query information is to the sub-tag corresponding to the fuzzy keyword of this category. The average value of the association probability represents the degree of association between the sub-tag corresponding to each fuzzy tag and the main tag. The greater the degree of association, the more similar the direction expressed by the query information is to the main tag corresponding to the fuzzy keyword of this category. Therefore, the present invention digitizes the degree of similarity between the fuzzy keywords of each category and the corresponding main tag and sub-tag by calculating the fuzzy probability coefficient of the fuzzy keywords of each category, and then uses it as a standard for judging the clarity and ambiguity of the intention of the query information, so as to reduce invalid push by verifying the push feedback information.

[0046] In particular, in the present invention, if the representation pointing category of the query information is clear representation pointing, the representation sample label corresponding to the query information is identified, and the feedback information associated with the main label is pushed preferentially; if the representation pointing category of the query information is fuzzy representation pointing, the query information is checked twice. In actual situations, the clear representation pointing indicates that the intention expressed by the query information is clear, so the feedback information associated with the main label corresponding to the standard keyword or the main label corresponding to the fuzzy keyword of the category corresponding to the probability consistency coefficient is pushed preferentially to the user end; the fuzzy representation pointing indicates that the intention expressed by the query information is unclear, that is, it indicates that the query information is unclear. The fuzzy keywords have poor representation of the intention. Therefore, by comparing the remaining keywords in the query information except the fuzzy keywords with the sample sentences in the sample database sentence by sentence, the main label most associated with the remaining keywords in the query information except the fuzzy keywords is determined, that is, the main label corresponding to the maximum value in the association representation coefficient. The feedback information associated with the most associated main label is pushed to the user end first. By adopting different verification methods for the query information with clear representation and fuzzy representation, the efficiency and effect of data verification are improved, and then the query information is fed back according to the verification results, and the probability of invalid feedback is reduced through data verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a schematic diagram of the steps of the information verification method based on big data according to an embodiment of the invention;

[0048] Figure 2 A flowchart of identifying keywords in query information according to an embodiment of the present invention;

[0049] Figure 3 A flowchart for determining whether the categories of fuzzy keywords in an embodiment of the invention are the same;

[0050] Figure 4 This is a flowchart for determining the category of query information representation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below with reference to embodiments. It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0052] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0053] See also Figures 1 to 4The following are a schematic diagram of the steps of the information verification method based on big data according to an embodiment of the present invention, a flowchart for determining whether keywords in query information are of the same category, a flowchart for determining whether the categories of fuzzy keywords are the same, and a flowchart for determining the category of the representation of the query information. The information verification method based on big data according to the present invention includes:

[0054] Step S1: Obtain a query message posted by a user on the interactive platform, compare the keywords in the query message with sample tags, and identify fuzzy keywords and standard keywords in the query message. The sample tags each include a main tag and several sub-tags, and both the sub-tags and the main tag are keywords.

[0055] Step S2: if the query information does not contain standard keywords, classify the fuzzy keywords according to the sub-tags corresponding to the fuzzy keywords, determine the fuzzy probability coefficients of the fuzzy keywords in each category, and determine the maximum fuzzy probability coefficient as the probability consistency coefficient;

[0056] Step S3, checking the representation direction category of the inquiry information according to the probability consistency coefficient, where the representation direction category includes fuzzy representation direction and clear representation direction;

[0057] Step S4, selecting a feedback method for the inquiry information according to the verification result, including:

[0058] Identify the sample characterization tag corresponding to the query information, and give priority to pushing feedback information associated with the sample characterization tag;

[0059] Alternatively, the query information is subjected to a secondary check, including identifying fuzzy keywords in each text sentence of the query information, determining the association representation coefficient of each main tag based on the comparison results of the remaining keywords other than the fuzzy keywords in each text sentence with each sample sentence in the sample database, and giving priority to pushing feedback information associated with the main tag corresponding to the largest association representation coefficient.

[0060] Specifically, the present invention does not specifically limit the setting method of feedback information. Preferably, in the prior art, feedback information often adopts a trigger mechanism, that is, when a certain keyword is triggered, a number of feedback information associated with the keyword is fed back. In this embodiment, the association relationship between the keyword and the number of feedback information can be pre-established;

[0061] For example, in a bank transaction platform, the keyword is loan, and some loan business explanation information is set as feedback information. Those skilled in the art can set it according to the needs of the application scenario, which will not be repeated here.

[0062] Specifically, the step S1 also includes pre-building sample labels, and the building process includes:

[0063] Obtain several sample sentences containing a single keyword from the sample database in advance, extract several remaining keywords from the sample sentences, and calculate the occurrence probability of each remaining keyword one by one.

[0064] If the occurrence probability of the remaining keywords is greater than a predetermined probability threshold, the single keyword is used as the main tag, the remaining keywords are used as subtags of the main tag, and the occurrence probability of the remaining keywords is determined as the association probability between the main tag and the subtag.

[0065] The probability of the remaining keywords appearing in the sample sentence is the probability of the remaining keywords appearing in the sample sentence.

[0066] Specifically, the present invention does not limit the specific configuration method of the sample database. The data therein may be obtained in advance by collecting query information from the user end in the interactive platform, or by other methods, which will not be described in detail.

[0067] Specifically, in this embodiment, the probability threshold is selected from the interval [0.15, 0.25].

[0068] Specifically, in the present invention, sample labels are pre-constructed. In actual situations, if a certain keyword appears in a sentence, the probability of the remaining keywords in the sentence appearing is high, indicating that this keyword and the remaining keywords are highly correlated in practice. Therefore, the keyword is used as the main label, and the remaining keywords with a high probability of appearing are used as sub-labels. By analyzing big data, the connection between the keywords can be reliably found, providing an important basis for the subsequent verification of the inquiry information.

[0069] For more details, please refer to Figure 2 As shown, in step S1, the process of identifying the fuzzy keywords and standard keywords in the query information based on the comparison results of the keywords in the query information and the sample labels includes:

[0070] If the keyword in the query information is the same as the main tag of the sample tag, the keyword is identified as a standard keyword;

[0071] If the keyword in the query information is the same as the sub-tag of the sample tag, the keyword is identified as an ambiguous keyword.

[0072] Specifically, in the present invention, fuzzy keywords and standard keywords in the query information are identified based on the comparison results of the keywords in the query information and the sample labels. In actual situations, if the keywords in the query information are the same as the main labels, it indicates that the intention of the query information is clearly expressed. If the keywords in the query information are the same as the sub-labels, it indicates that the intention of the query information may not be clear and needs further verification. Therefore, the keywords that are the same as the main labels are determined as standard keywords, and the keywords that are the same as the sub-labels are determined as fuzzy keywords, so as to distinguish the keywords in the query information, further verify whether the intention expressed by the query information is clear, and feedback is given to the query information based on the verification results, thereby reducing the probability of invalid feedback through data verification.

[0073] For more details, please refer to Figure 3 As shown, in step S2, the process of classifying fuzzy keywords according to the sub-tags corresponding to the fuzzy keywords includes:

[0074] Based on the sub-tags corresponding to each fuzzy keyword, determine whether the categories of each fuzzy keyword are the same.

[0075] If the subtags corresponding to the fuzzy keywords belong to the same main tag, it is determined that the categories of the fuzzy keywords are the same.

[0076] Specifically, in the present invention, fuzzy keywords are classified according to the sub-tags corresponding to the fuzzy keywords. In actual situations, if the main tags to which the sub-tags corresponding to the fuzzy keywords belong are the same, it indicates that the intentions expressed by the fuzzy keywords are the same. Therefore, the fuzzy keywords are regarded as the same category, which facilitates the subsequent verification of whether the intentions expressed in the query information are clear based on the situation of the fuzzy keywords in each category, so as to provide effective feedback on the query information through the verification results.

[0077] Specifically, in step S2, the fuzzy probability coefficient of each category of fuzzy keywords is determined, where:

[0078] According to formula (1), the fuzzy probability coefficient Ki corresponding to the fuzzy keyword belonging to category i is calculated.

[0079] ,

[0080] In formula (1), N0 represents the total number of fuzzy keywords in the query information, n represents the number of fuzzy keywords belonging to the i-th category, Pa represents the association probability between the sub-tag corresponding to the a-th fuzzy keyword belonging to the i-th category and the associated main tag, and i and a are both integers greater than 0.

[0081] Specifically, in the present invention, the fuzzy probability coefficient of each category of fuzzy keywords is determined. The fuzzy probability coefficient is calculated by the proportion of fuzzy keywords belonging to the same category in the fuzzy keywords in the query information and the average value of the association probability between the sub-tags corresponding to each fuzzy keyword belonging to this category and the associated main tag. In actual situations, the larger the proportion, the more similar the direction expressed by the query information is to the sub-tag corresponding to the fuzzy keyword of this category. The average value of the association probability represents the degree of association between the sub-tag corresponding to each fuzzy tag and the main tag. The greater the degree of association, the more similar the direction expressed by the query information is to the main tag corresponding to the fuzzy keyword of this category. Therefore, the present invention digitizes the degree of similarity between the fuzzy keywords of each category and the corresponding main tag and sub-tag by calculating the fuzzy probability coefficient of the fuzzy keywords of each category, and then uses it as a standard for judging the clarity and ambiguity of the intention of the query information, so as to reduce invalid push by verifying the push feedback information.

[0082] For more details, please refer to Figure 4 As shown, in step S3, the process of checking the representation pointing category of the query information according to the probability consistency coefficient includes:

[0083] Compare the probability consistency coefficient Km with the preset coefficient comparison threshold K0,

[0084] If Km≥Km0, verify and determine that the representation direction category of the query information is clear representation direction;

[0085] If Km<Km0, it is checked to determine that the representation orientation category of the query information is fuzzy representation orientation.

[0086] Specifically, in this embodiment, if the average value △P of the association probability between the sub-tags corresponding to all fuzzy keywords in the query information and the associated main tags is less than 0.5, K0 is selected from the interval [0.15, 0.25]; if the average value △P of the association probability between the sub-tags corresponding to all fuzzy keywords in the query information and the associated main tags is greater than or equal to 0.5, K0 is selected from the interval [0.25, 0.4].

[0087] Specifically, the step S3 further includes determining the representation direction category of the query information based on the keywords in the query information,

[0088] If the query information contains standard keywords, it is determined that the representation orientation category of the query information is clear representation orientation.

[0089] Specifically, in step S4, the process of selecting a feedback method for the inquiry information based on the verification result includes:

[0090] If the representation direction category of the query information is clear representation direction, identify the representation sample tag corresponding to the query information, and give priority to pushing feedback information associated with the main tag;

[0091] If the representation direction category of the query information is fuzzy representation direction, a secondary check is performed on the query information.

[0092] Specifically, in step S4, the process of identifying the sample label corresponding to the query information includes:

[0093] If the query information contains a standard keyword, the main tag corresponding to the standard keyword is determined as the representative sample tag;

[0094] If there is no standard keyword in the query information, a target category fuzzy keyword is determined, and the main label of the sub-label corresponding to the target category fuzzy keyword is used as the representative sample label. The target category fuzzy keyword is a fuzzy keyword of the category corresponding to the probability consistency coefficient.

[0095] Specifically, in step S4, the process of determining the association representation coefficient of each main tag based on the comparison results of the remaining keywords excluding the fuzzy keywords in each text sentence with each sample sentence in the sample database includes:

[0096] The occurrence probability of each of the remaining keywords in the sample sentence containing a single sub-tag is determined, and an average value of the occurrence probabilities is determined as the association representation coefficient of the main tag.

[0097] In this embodiment, the occurrence probability of a single remaining keyword is set to A=An / An0, where An0 represents the number of sample sentences containing a single sub-tag, and An0 represents the number of sample sentences containing a single remaining keyword.

[0098] Specifically, in the present invention, if the representation pointing category of the query information is clear representation pointing, the representation sample label corresponding to the query information is identified, and the feedback information associated with the main label is pushed first. If the representation pointing category of the query information is fuzzy representation pointing, the query information is checked twice. In actual situations, the clear representation pointing indicates that the intention expressed by the query information is clear, so the feedback information associated with the main label corresponding to the standard keyword or the main label corresponding to the fuzzy keyword of the category corresponding to the probability consistency coefficient is pushed to the user end first. The fuzzy representation pointing indicates that the intention expressed by the query information is unclear, which means that the query information is unclear. The fuzzy keywords in the query information have poor representation of the intent. Therefore, by comparing the remaining keywords in the query information except the fuzzy keywords with the sample sentences in the sample database sentence by sentence, the main label most associated with the remaining keywords in the query information except the fuzzy keywords is determined, that is, the main label corresponding to the maximum value in the association representation coefficient. The feedback information associated with the most associated main label is pushed to the user end first. By adopting different verification methods for query information with clear representation and fuzzy representation, the efficiency and effect of data verification are improved, and then feedback is given to the query information based on the verification results, and the probability of invalid feedback is reduced through data verification.

[0099] If the big data-based information verification method of the present invention is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention, and the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0100] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. A method for information verification based on big data, characterized in that: include: Step S1: Obtain a query message posted by a user on the interactive platform, compare the keywords in the query message with sample tags, and identify fuzzy keywords and standard keywords in the query message. The sample tags each include a main tag and several sub-tags, and both the sub-tags and the main tag are keywords. Step S2: if the query information does not contain standard keywords, classify the fuzzy keywords according to the sub-tags corresponding to the fuzzy keywords, determine the fuzzy probability coefficients of the fuzzy keywords in each category, and determine the maximum fuzzy probability coefficient as the probability consistency coefficient; Step S3, checking the representation direction category of the inquiry information according to the probability consistency coefficient, where the representation direction category includes fuzzy representation direction and clear representation direction; Step S4, selecting a feedback method for the inquiry information according to the verification result, including: If the representation direction category of the query information is clear representation direction, identifying the representation sample tag corresponding to the query information, and preferentially pushing feedback information associated with the representation sample tag; If the representation orientation category of the query information is fuzzy representation orientation, a secondary check is performed on the query information, including identifying fuzzy keywords in each text sentence of the query information, determining the relevance representation coefficient of each main tag based on the comparison results of the remaining keywords in each text sentence excluding the fuzzy keywords with each sample sentence in the sample database, and preferentially sending feedback information associated with the main tag corresponding to the largest relevance representation coefficient; In step S1, the process of identifying fuzzy keywords and standard keywords in the query information based on the comparison results of the keywords in the query information and the sample labels includes: If the keyword in the query information is the same as the main tag of the sample tag, the keyword is identified as a standard keyword; If the keyword in the query information is the same as the sub-tag of the sample tag, the keyword is identified as a fuzzy keyword; In step S2, the fuzzy probability coefficient of each category of fuzzy keywords is determined, wherein, According to formula (1), the fuzzy probability coefficient Ki corresponding to the fuzzy keyword belonging to the i-th category is calculated. In formula (1), N0 represents the total number of fuzzy keywords in the query information, n represents the number of fuzzy keywords belonging to the i-th category, Pa represents the association probability between the sub-tag corresponding to the a-th fuzzy keyword belonging to the i-th category and the associated main tag, and i and a are both integers greater than 0; In step S3, the process of checking the representation pointing category of the query information according to the probability consistency coefficient includes: Compare the probability consistency coefficient with a preset coefficient comparison threshold, If the probability consistency coefficient is greater than or equal to the coefficient comparison threshold, verify and determine that the representation direction category of the inquiry information is clear representation direction; If the probability consistency coefficient is less than the coefficient comparison threshold, verify and determine that the representation direction category of the query information is fuzzy representation direction; In step S4, the process of determining the correlation representation coefficient of each main tag based on the comparison results of the remaining keywords excluding the fuzzy keywords in each text sentence with each sample sentence in the sample database includes: The occurrence probability of each of the remaining keywords in the sample sentence containing a single sub-tag is determined, and an average value of the occurrence probabilities is determined as the association representation coefficient of the main tag.

2. The information verification method based on big data according to claim 1, characterized in that: The step S1 also includes pre-building sample labels, and the building process includes: Obtain several sample sentences containing a single keyword from the sample database in advance, extract several remaining keywords from the sample sentences, and calculate the occurrence probability of each remaining keyword one by one. If the probability of occurrence of the remaining keywords is greater than the predetermined probability threshold, the single keyword is used as the main tag, the remaining keywords are used as sub-tags of the main tag, and the probability of occurrence of the remaining keywords is determined as the association probability between the main tag and the sub-tag. The probability of occurrence of the remaining keywords is the probability of the remaining keywords appearing in the sample sentence.

3. The information verification method based on big data according to claim 1, characterized in that: In step S2, the process of classifying the fuzzy keywords according to the sub-tags corresponding to the fuzzy keywords includes: Based on the sub-tags corresponding to each fuzzy keyword, determine whether the categories of each fuzzy keyword are the same. If the subtags corresponding to the fuzzy keywords belong to the same main tag, it is determined that the categories of the fuzzy keywords are the same.

4. The information verification method based on big data according to claim 1, characterized in that: The step S3 further includes determining the representation direction category of the query information based on the keywords in the query information, If the query information contains standard keywords, it is determined that the representation orientation category of the query information is clear representation orientation.

5. The information verification method based on big data according to claim 1, characterized in that: In step S4, the process of identifying the sample label corresponding to the query information includes: If the query information contains a standard keyword, the main tag corresponding to the standard keyword is determined as the representative sample tag; If there is no standard keyword in the query information, a target category fuzzy keyword is determined, and the main label of the sub-label corresponding to the target category fuzzy keyword is used as the representative sample label. The target category fuzzy keyword is a fuzzy keyword of the category corresponding to the probability consistency coefficient.

Citation Information

Patent Citations

  • Enterprise talent information management system based on big data

    CN111445212A

  • Search intention identification method and device

    CN105095187A

  • Speech emotion recognition method based on data analysis

    CN116884392A