Key industry information industry map construction method, system and terminal
By extracting and sorting keyword information in industry questionnaires and building an industry map, the problem of low efficiency in obtaining industry information in the existing technology is solved, and faster and more accurate information acquisition is achieved.
Patent Information
- Application Number
- CN202510321700.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-17
AI Technical Summary
Existing search engines are inefficient when enterprises obtain industry information, have messy information, less targeted information, and limited information scope.
By obtaining industry questionnaire information, extracting keywords and giving priority weights, combining them to form keyword combinations, inputting industry matching database to determine the initial inspection industry information, sorting key industry information, and building an industry map based on the correlation threshold.
It improves the efficiency of enterprises to obtain relevant industry information, and users can obtain the required information more quickly through the industry map, improving market response speed and decision-making capabilities.
Smart Images

Figure CN120163221A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of atlas construction, and particularly to a method, system and terminal for constructing an industry atlas of key industry information. Background Art
[0002] An industry atlas refers to a structured description of various business entities, concepts, attributes and relationships in a specific industry business field, which presents industry knowledge in a graphical way, thereby helping enterprises quickly master the market, products and related information, and improving the market response speed and decision-making ability.
[0003] Currently, enterprises usually adopt the method of network query when expanding business and grasping the market trend. When querying information, relevant industry information is obtained by inputting information retrieval terms of the industry. However, the information obtained by searching through existing search engines is relatively messy, with less targeted information, and the scope of industry information that can be retrieved is limited, resulting in low efficiency. Summary of the Invention
[0004] In order to improve the efficiency of enterprises in obtaining relevant industry information of their own industry, the present invention provides a method, system and terminal for constructing an industry atlas of key industry information.
[0005] In a first aspect, the present invention provides a method for constructing an industry atlas of key industry information, adopting the following technical solutions:
[0006] A method for constructing an industry atlas of key industry information includes:
[0007] Obtaining industry questionnaire information by a preset information collection method;
[0008] Inputting the industry questionnaire information into a preset keyword library to extract all keyword information in the industry questionnaire information, and assigning a priority weight to each keyword information;
[0009] Combining single or multiple keyword information to form keyword combinations, and determining the total priority weight of the keyword combinations;
[0010] Inputting each keyword combination into a preset industry matching database to determine the preliminary inspection industry information corresponding to the keyword combination;
[0011] Sorting the preliminary inspection industry information according to the total priority weight to obtain the preliminary inspection industry information corresponding to the maximum total priority weight and defining it as key industry information;
[0012] Determining the upstream and downstream position information according to the key industry information and a preset relevance industry threshold;
[0013] Construct an industry map based on key industry information, upstream and downstream location information, and a preset industry database.
[0014] By adopting the above technical solution, the system obtains industry questionnaire information and identifies and analyzes the content keywords therein to obtain key industry information, then determines its upstream and downstream information and the position of the key industry information in the upstream and downstream information, and finally constructs an industry map based on this. After establishing the industry map, users can obtain the industry information required by this industry more efficiently through the industry map.
[0015] Optionally, it further includes:
[0016] Eliminate the key industry information in the initially inspected industry information, and define the remaining initially inspected industry information as auxiliary industry information;
[0017] Re - sort the auxiliary industry information in descending order according to the total priority weight to determine the matching order;
[0018] Perform a similarity detection on the auxiliary industry information and the key industry information according to the matching order and a preset relevant industrial cluster to obtain the industry similarity;
[0019] Based on the industry similarity being greater than a preset similarity threshold, determine the upstream and downstream position information of the auxiliary industry according to the auxiliary industry information and the relevant industry threshold;
[0020] Construct a relevant industry map according to the auxiliary industry information, the auxiliary upstream and downstream position information, and the industry database;
[0021] Merge the industry map and the relevant industry map to obtain a new industry map.
[0022] Optionally, invalid words need to be eliminated before combining keyword information. Invalid words include duplicate words and synonymous words. The methods for processing invalid words include:
[0023] Classify all keyword information to determine the frequency of occurrence of each keyword information;
[0024] When there is keyword information with a frequency of occurrence greater than 1, define the keyword information with a frequency of occurrence greater than 1 as a duplicate word, and eliminate the redundant duplicate words, leaving only 1;
[0025] After elimination, input pairwise keyword information into a preset synonymous word library for matching to determine whether they are synonymous words;
[0026] When pairwise keyword information are synonymous words, compare the frequencies of occurrence of the pairwise keyword information before elimination;
[0027] The keyword information with high frequency of occurrence is retained, and the keyword information with low frequency of occurrence is eliminated.
[0028] Optionally, the information collection method includes:
[0029] Obtaining a behavior image input by a person;
[0030] Determine the input behavior type according to the input behavior image of the person and the preset input tool characteristics, the input behavior type includes web questionnaire input and paper questionnaire input;
[0031] Based on the webpage questionnaire input, the completed webpage questionnaire content is identified by a preset webpage questionnaire content identification method to obtain the industry questionnaire information and output it;
[0032] Based on the paper questionnaire input, the completed paper questionnaire content is identified using a preset paper questionnaire content recognition method to obtain industry questionnaire information and output it.
[0033] Optionally, the webpage questionnaire content identification method includes:
[0034] Obtaining character information and character input speed within a preset input recognition area;
[0035] When the character input speed is greater than the preset normal input speed, the character information is input into a preset character database for analysis to determine whether the character information is recognizable;
[0036] When the character information is identifiable, the character information is defined as valid input information, otherwise it is defined as invalid input information, and the invalid input information includes blank information, non-text character information and incoherent sentence information;
[0037] Determine the effective content ratio of the webpage questionnaire content according to the invalid input content and the effective input content;
[0038] When the proportion of valid content is not greater than the preset invalidation ratio, the webpage questionnaire content will be invalidated;
[0039] When the proportion of valid content is greater than the invalid proportion, all valid input information in the web page questionnaire content will be combined to form industry questionnaire information and then output.
[0040] Optionally, also include:
[0041] When the character input speed is not greater than the normal input speed, determining whether there is invalid input information according to the character information and the character database;
[0042] If there is invalid input information, obtain the input state of the personnel when filling in the web questionnaire;
[0043] When and only when the facial expression of the personnel input is consistent with the preset serious input facial expression, determine the proportion of the valid content of the web questionnaire content according to the invalid input content and the valid input content;
[0044] Based on the proportion of valid content being no greater than the invalidation ratio, obtain the facial expression of the personnel input when the preset submission button is triggered;
[0045] When and only when the facial expression of the personnel input is consistent with the preset panicked input facial expression, merge all the valid input information of the completed part in the web questionnaire content to form industry questionnaire information and then output it;
[0046] Based on the proportion of valid content being greater than the invalidation ratio, input the invalid input information into the preset word correction library for identification and analysis to obtain corrected input information;
[0047] Merge the corrected input information and the valid input information to form industry questionnaire information and then output it.
[0048] Optionally, the method for identifying the content of a paper questionnaire includes:
[0049] Obtain the image information of the paper questionnaire;
[0050] According to the image information of the paper questionnaire and the preset questionnaire text specification, frame out all the writing areas in the image information of the paper questionnaire;
[0051] Number all the writing areas to obtain the recognition order;
[0052] According to the recognition order, extract the written content of the personnel in the writing area according to the preset character recognition method;
[0053] When and only when the written content of the personnel is inconsistent with the preset content format of the writing area, increase the invalidation risk value by one, and accumulate it in sequence to obtain the cumulative risk value;
[0054] When the cumulative risk value is greater than the preset benchmark risk value, invalidate the paper questionnaire;
[0055] When the cumulative risk value is not greater than the benchmark risk value, merge the written content of the personnel to form industry questionnaire information and then output it.
[0056] Optionally, the processing method when the writing area is covered by stains includes:
[0057] Determine the stain position according to the image information of the paper questionnaire and the preset stain characteristics;
[0058] Determine the contaminated character content in the written content of the personnel contaminated by stains according to the stain position and the writing area;
[0059] The polluted characters are removed from the written content, and the remaining written content is input into a preset word correction library for correction to obtain a corrected written text;
[0060] Input the corrected written text and the written content of people in other writing areas into a preset correlation database for correlation matching to obtain a correlation value;
[0061] When the correlation value is greater than the preset reference correlation, the corrected writing text is combined with the writing content of people in other writing areas to form industry questionnaire information and then output;
[0062] When the correlation value is not greater than the reference correlation value, the corrected written text is discarded, and a preset transparent liquid is dripped into the stain position, and the handwriting path segment is obtained through the spherical transparent liquid within a preset penetration time;
[0063] Input the handwriting path fragments into a preset character restoration library for restoration to obtain restored writing text;
[0064] The restored written text and the uncontaminated written content of personnel are combined to form industry questionnaire information and then output.
[0065] In the second aspect, the present application provides a key industry information industry map construction system, which adopts the following technical solutions:
[0066] A key industry information industry graph construction system, applying any of the above key industry information industry graph construction methods, includes:
[0067] An acquisition module is used to acquire the input behavior image of the personnel, character information, character input speed, the input demeanor of the personnel when filling in the web questionnaire, the input demeanor of the personnel when the preset submit button is triggered, the image information of the paper questionnaire, and the handwriting path fragment;
[0068] A memory for storing a program for a method for constructing a key industry information industry map;
[0069] The program in the processor and the memory can be loaded and executed by the processor and realize a method for constructing a key industry information industry map.
[0070] In a third aspect, the present application provides a smart terminal, which adopts the following technical solution:
[0071] An intelligent terminal includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute any of the above-mentioned key industry information industry map construction methods.
[0072] In summary, the present application includes at least one of the following beneficial technical effects:
[0073] 1. The system obtains industry questionnaire information and identifies and analyzes the content keywords therein to obtain key industry information. Then, based on the key industry information, it determines its upstream and downstream information and the position of the key industry information in the upstream and downstream information, and finally constructs an industry map accordingly. After establishing the industry map, users can obtain the industry information required for their industry more efficiently through the industry map;
[0074] 2. By obtaining the expressions of web questionnaire fillers when they enter invalid input information, if the person is in a serious input expression when filling in the information and in a panicked input expression when triggering the submission button, it indicates that the person's filling of invalid input information is an inadvertent mistake, and the system will handle the above situation as appropriate;
[0075] 3. When the content of a paper questionnaire is contaminated by stain features, the system first corrects the uncontaminated part of the person's writing content. If the corrected content has a high relevance to the context content, it can be directly used; otherwise, a transparent liquid is dropped at the stain position, and the text is restored by obtaining the handwriting path segments through the transparent liquid. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a flowchart of a method for constructing an industry map of key industry information according to an embodiment of the present invention;
[0077] Figure 2 is a flowchart of a method for processing auxiliary industry information other than key industry information in the preliminary inspection of industry information according to an embodiment of the present invention;
[0078] Figure 3 is a flowchart of a method for processing invalid words according to an embodiment of the present invention;
[0079] Figure 4 is a flowchart of an information collection method according to an embodiment of the present invention;
[0080] Figure 5 is the flowchart of a method for identifying the content of a web questionnaire according to an embodiment of the present invention Figure 1 ;
[0081] Figure 6 is the flowchart of a method for identifying the content of a web questionnaire according to an embodiment of the present invention Figure 2 ;
[0082] Figure 7 is the flowchart of a method for identifying the content of a paper questionnaire according to an embodiment of the present invention Figure 1 ;
[0083] Figure 8 is the flowchart of a method for identifying the content of a paper questionnaire according to an embodiment of the present invention Figure 2 . DETAILED DESCRIPTION OF THE INVENTION
[0084] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0085] An embodiment of the present application discloses a method for constructing an industry map of key industry information. The system obtains industry information through web questionnaires and paper questionnaires, and identifies and analyzes the collected information. After obtaining relevant industry information content, it is saved in the system database to summarize and form an industry map.
[0086] Refer to Figure 1 , a method for constructing an industry map of key industry information includes the following steps:
[0087] Step S100: Obtain industry questionnaire information by a preset information collection method.
[0088] Industry questionnaire information refers to data information containing content of various industries obtained through external collection. Industry questionnaire information is obtained by an information collection method. The information collection method mainly collects web file information and paper questionnaire information. The information collection method will not be elaborated here and will be introduced in detail in subsequent embodiments.
[0089] Step S101: Input the industry questionnaire information into a preset keyword library to extract all keyword information in the industry questionnaire information, and assign a priority weight to each keyword information.
[0090] The keyword library is a database tool designed by technicians for extracting keywords in text information, which will not be elaborated here.
[0091] Keyword information refers to nominal keywords in the text of industry questionnaire information that can be used to determine relevant situations in this industry. A large number of nominal keywords are included in the industry questionnaire information formed when users fill out questionnaires. The industry information of the users can be inferred from these nominal keywords. By inputting the industry questionnaire information into the keyword library, all keyword information can be extracted from the industry questionnaire information.
[0092] The priority weight is used as a mark when inferring the industry information corresponding to the industry questionnaire information later. In this embodiment, the priority weight is set to 1, that is, the priority weight of each keyword information is 1.
[0093] Step S102: Combine single or multiple keyword information to form keyword combinations, and determine the total priority weight of the keyword combinations.
[0094] When analyzing the text, different combinations of keyword information can infer different industry information, and the more keywords are combined, the closer the industry information is to the industry information that the person filling out the questionnaire wants to express. By adding up the priority weights of all keyword information, the total priority weight of the keyword combination can be determined. The higher the total priority weight of the keyword combination, the more accurate the industry information inferred.
[0095] Step S103: input each keyword combination into a preset industry matching database to determine the initial inspection industry information corresponding to the keyword combination.
[0096] The industry matching database is a database tool set up by technical personnel to analyze keyword combinations and determine corresponding industry information, which will not be described in detail here.
[0097] Initial inspection industry information refers to the industry information inferred by analyzing each keyword combination. Different keyword combinations can analyze different initial inspection industry information.
[0098] Step S104: sort the initial inspection industry information according to the sum of priority weights to obtain the initial inspection industry information corresponding to the maximum sum of priority weights and define it as key industry information.
[0099] Key industry information refers to the industry information that users who fill out the questionnaire want to retrieve. After sorting the initial inspection industry information according to the total priority weight, the initial inspection industry information with the largest total priority weight uses the most keywords for combination. The initial inspection industry information is closer to the industry information that users want to retrieve, that is, the key industry information.
[0100] Step S105: Determine upstream and downstream location information based on key industry information and preset correlation industry thresholds.
[0101] The relevant industry threshold is set by technical personnel to judge the upstream and downstream conditions of the industry. It is determined by analyzing the upstream and downstream conditions of key industry information and will not be elaborated here.
[0102] Upstream and downstream position information refers to the position of key industry information in the upstream and downstream of the industrial chain. Only by clarifying the position of the industry in the upstream and downstream of the industrial chain can enterprises or users search for the required industry information from the upstream or downstream industries in a targeted manner.
[0103] After the key industry information is determined from the industry questionnaire information, the key industry information is analyzed through the relevant industry threshold to determine the upstream and downstream position information of the key industry information.
[0104] Step S106: Construct an industry map based on key industry information, upstream and downstream location information, and a preset industry database.
[0105] The industry database is a database of industry information that has been saved in the system database. It is a database formed by the system after long-term collection of external data. It contains the recorded upstream and downstream conditions of various industries, which will not be elaborated here.
[0106] After the upstream and downstream location information of the key industry information is determined, the upstream and downstream location information of the key industry information is summarized with the existing industry database, and the key industry information can be entered into the industry database to form an industry map.
[0107] Reference Figure 2 The method for processing auxiliary industry information other than key industry information in the initial inspection industry information includes the following steps:
[0108] Step S200: remove key industry information from the initial inspection industry information, and define the remaining initial inspection industry information as auxiliary industry information.
[0109] The key industry information in the initial inspection industry information is the industry indicated by the user filling out the questionnaire, while other industry information in the initial inspection industry information, namely the auxiliary industry information, may be industries related to the key industry information. This industry information can play an auxiliary role in constructing the key industry map and broaden the upstream and downstream information of the key industry.
[0110] Step S201: Rearrange the auxiliary industry information in descending order according to the sum of priority weights to determine a matching order.
[0111] After removing the priority weight sum of the key industry information, the larger the priority weight sum, the stronger the correlation between the auxiliary industry information and the key industry information, and similarly to step S104, each auxiliary industry information is analyzed in matching order after sorting. The matching order refers to the order in which each auxiliary industry information is analyzed.
[0112] Step S202: Perform similarity detection on the auxiliary industry information and the key industry information according to the matching order and the preset related industry clusters to obtain industry similarity.
[0113] The related industry cluster is a tool set up for technicians to determine the similarity between auxiliary industry information and key industry information, which will not be elaborated here.
[0114] Industry similarity refers to the similarities between two industries, mainly the upstream and downstream industrial chains. The higher the degree of consistency between the upstream and downstream industrial chains, the higher the similarity.
[0115] The auxiliary industry information and key industry information are analyzed to determine if the industry similarity between the two is high enough, and only then can the auxiliary industry information be used to assist in the construction of the industry map of the key industry information.
[0116] Step S203: Based on the industry similarity being greater than a preset similarity threshold, determine the upstream and downstream position information of the auxiliary industry according to the auxiliary industry information and the relevance industry threshold.
[0117] The similarity threshold is a standard set by technicians to measure whether two industries are related industries. Only when the industry similarity between the two is greater than the similarity threshold can it be shown that the two are related industries, which will not be elaborated here.
[0118] When the industry similarity is not greater than the similarity threshold, it indicates that the similarity between the auxiliary industry information and the key industry information is not high, and this auxiliary industry information cannot be used to assist in constructing the industry map of the key industry information.
[0119] When the industry similarity is greater than the similarity threshold, it indicates that the auxiliary industry information and the key industry information are related industries. At this time, the auxiliary industry information can be used to assist in constructing the industry map of the key industry information.
[0120] Similar to step S105, determine the upstream and downstream position information of the auxiliary industry of the auxiliary industry information.
[0121] Step S204: Construct a related industry map according to the auxiliary industry information, the auxiliary upstream and downstream position information, and the industry database.
[0122] Similar to step S106, construct a related industry map of the auxiliary industry information according to the auxiliary industry information.
[0123] Step S205: Merge the industry map and the related industry map to obtain a new industry map.
[0124] Since the auxiliary industry information and the key industry information are related industries, there are commonalities between the industry map of the key industry information and the related industry map of the auxiliary industry information. Merge the two to form a new industry map, which contains the upstream and downstream information of both the auxiliary industry information and the key industry information. When users refer to the upstream and downstream information of the key industry information, they can also refer to the upstream and downstream information of its related auxiliary industry information.
[0125] Refer to Figure 3 , invalid words need to be removed before combining keyword information. Invalid words include duplicate words and synonymous words. The methods for processing invalid words include the following steps:
[0126] Step S300: Classify all keyword information to determine the frequency of occurrence of each keyword information.
[0127] The frequency of a term's occurrence refers to the number of times the keyword information appears in the industry questionnaire information. Repeated keyword information will affect the calculation of the priority weight, and the repeated terms will be removed according to the frequency of term occurrence.
[0128] Step S301: When there is keyword information with a frequency of occurrence greater than 1, define the keyword information with a frequency of occurrence greater than 1 as repeated terms, and remove the redundant repeated terms, leaving only 1.
[0129] The keyword information with a frequency of occurrence not greater than 1 appears only once and does not need to be removed. If there is keyword information with a frequency of occurrence greater than 1, then this keyword information is a repeated term. Only one keyword information needs to be left among the keyword information extracted from the industry questionnaire information, and the rest are all removed.
[0130] Step S302: After the removal, input two pieces of keyword information into a preset thesaurus of near-synonyms to determine whether they are near-synonyms.
[0131] In this embodiment, there are still words with similar meanings among the keyword information. Near-synonyms will also affect the calculation of the priority weight, and the near-synonyms need to be removed.
[0132] The thesaurus of near-synonyms is a database set by technicians to determine whether two words are near-synonyms, which will not be elaborated here. By inputting two words in the keyword information into the thesaurus of near-synonyms, it can be determined whether the two are near-synonyms.
[0133] Step S303: When two pieces of keyword information are near-synonyms, compare the frequencies of occurrence of the two pieces of keyword information before removal.
[0134] If the two are not near-synonyms, continue to compare other keyword terms until all keyword terms have been compared.
[0135] If the two are near-synonyms, then one of the words needs to be removed. The method of removal is to compare the frequencies of occurrence of the two words in the industry questionnaire information. The higher the frequency of occurrence, the higher the reference value.
[0136] Step S304: Retain the keyword information with a high frequency of occurrence and remove the keyword information with a low frequency of occurrence.
[0137] After determining two keyword terms, retain the keyword information with a high frequency of occurrence. If the frequencies of occurrence of the two keyword terms are the same, then choose one to remove.
[0138] Refer to Figure 4 , the information collection method includes the following steps:
[0139] Step S400: Obtain the image of the person's input behavior.
[0140] In this embodiment, the information text is collected in the form of web questionnaires and paper questionnaires. Whether it is a web questionnaire or a paper questionnaire, the questionnaire fillers need to fill in the questionnaire. The image of the person's input behavior refers to the image of the questionnaire filler when filling in the questionnaire collected by the camera.
[0141] Step S401: Determine the input behavior type according to the image of the person's input behavior and the preset input tool features. The input behavior types include web questionnaire input and paper questionnaire input.
[0142] The input tool features refer to the features of computer devices or pens and papers.
[0143] By identifying and confirming the input tool features in the image of the person's input behavior, the input behavior type of the questionnaire filler can be determined. The input behavior type refers to whether the current questionnaire filler is filling in a web document or a paper questionnaire. When the computer device recognized by the camera, it indicates that the web questionnaire is being filled. If the pen and paper features are recognized, it indicates that the paper questionnaire is being filled.
[0144] Step S402: Based on the web questionnaire input, use the preset web questionnaire content recognition method to recognize the completed web questionnaire content to obtain the industry questionnaire information and output it.
[0145] When filling in the web questionnaire, the completed web questionnaire content is recognized by the web questionnaire content recognition method. The web questionnaire content recognition method will not be elaborated here and will be introduced in detail in the subsequent embodiments.
[0146] Step S403: Based on the paper questionnaire input, use the preset paper questionnaire content recognition method to recognize the completed paper questionnaire content to obtain the industry questionnaire information and output it.
[0147] For the paper questionnaire, the system recognizes the content in the paper questionnaire through the paper questionnaire content recognition method. The paper questionnaire content recognition method will not be elaborated here and will be introduced in detail in the subsequent embodiments.
[0148] Refer to Figure 5 , the web questionnaire content recognition method includes the following steps:
[0149] Step S500: Obtain the character information and character input speed within the preset input recognition area.
[0150] The input recognition area is the area position in the web questionnaire where the filler needs to fill in the information.
[0151] Character information refers to the information content filled in by the person filling out the web questionnaire. The character input speed refers to the typing speed of the questionnaire filler when inputting character information.
[0152] Both the character information and the character input speed are directly obtained by the system when the person filling out the information is filling in the information.
[0153] Step S501: When the character input speed is greater than the preset normal input speed, input the character information into the preset character database for analysis to determine whether the character information is recognizable.
[0154] The normal input speed is a reference value of the normal typing speed set by technicians for a person when seriously filling in web information. If the typing speed exceeds the normal input speed, it can be determined that there is an abnormality in the questionnaire filler when filling in the information, and there is a situation of random input, which will not be elaborated here.
[0155] The character database is a database of normal Chinese character texts that can be read continuously and coherently set by technicians, and is used to detect whether the text data is normal and can be read continuously and coherently, which will not be elaborated here.
[0156] Processing and analyzing the character information input by the questionnaire filler through the character database can determine whether the character information input by the filler is a normal Chinese character sentence that can be recognized and read continuously and coherently.
[0157] Step S502: When the character information is recognizable, define the character information as valid input information, otherwise define it as invalid input information. Invalid input information includes blank information, non-text character information, and incoherent sentence information.
[0158] Classify different types of recognizable character information. The recognizable ones are defined as valid input information, and the unrecognizable ones are defined as invalid input information. The system uses different methods to process different types of character information.
[0159] Step S503: Determine the proportion of valid content in the web questionnaire content based on the invalid input content and the valid input content.
[0160] The proportion of valid content refers to the proportion of valid input content in the character information filled in by the questionnaire filler. When both the invalid input content and the valid input content in the character information are confirmed, the proportion of valid content of the valid input content can be directly calculated.
[0161] Step S5041: When the proportion of valid content is not greater than the preset invalidation ratio, invalidate the web questionnaire content.
[0162] The invalidation ratio is a standard set by technicians to determine whether the web questionnaire content for normal recognition and analysis needs to be invalidated, which will not be elaborated here.
[0163] When the proportion of valid content is not greater than the invalid content proportion, it means that there are fewer valid input contents in the web questionnaire and more invalid input contents. The web questionnaire content is deemed to be filled out randomly by the person filling out the questionnaire, so the web questionnaire content needs to be invalidated.
[0164] Step S5042: When the effective content ratio is greater than the invalid content ratio, all the effective input information in the webpage questionnaire content is combined to form the industry questionnaire information and then output.
[0165] When the proportion of valid content is greater than the invalid content proportion, it means that there are more valid input contents in the web page questionnaire content, and fewer invalid input contents. The web page questionnaire content has reference value, and the invalid input information in the web page questionnaire content is eliminated before output.
[0166] Reference Figure 6 , when the character input speed is not greater than the normal input speed, the webpage questionnaire content recognition method comprises the following steps:
[0167] Step S600: When the character input speed is not greater than the normal input speed, determining whether there is invalid input information based on the character information and the character database.
[0168] When the character input speed is not greater than the normal input speed, it means that the person filling out the questionnaire is slow and is likely to be answering the questions seriously. If invalid input information appears at this time, it needs to be handled according to the situation and will not be directly invalidated. In order to ensure that there is no problem with the information filled out by the questionnaire filler, it is still necessary to first determine whether there is invalid input information.
[0169] Step S601: If there is invalid input information, the input state of the person when filling in the web questionnaire is obtained.
[0170] The facial expression of the person input refers to the facial expression of the person filling out the questionnaire when filling out the questionnaire. The facial expression of the person input is obtained by calling the camera of the computer device.
[0171] If there is no invalid input information, the content filled in by the person will be output directly.
[0172] If there is invalid input information, the system obtains the person's facial expression and determines how to handle the invalid input information based on the facial expression.
[0173] Step S602: If and only if the input attitude of the person is consistent with the preset serious input attitude, determine the effective content ratio of the webpage questionnaire content according to the invalid input content and the valid input content.
[0174] The serious input expression is a database obtained through training by a neural network model. This database includes images of the facial expression details of people when they are serious and focused, which will not be elaborated here.
[0175] Compare the input expression situation of the person with the serious input expression. If the two are inconsistent, it means that the person is not in a serious input state when filling in invalid input information. At this time, the invalid input information will be invalidated.
[0176] If the two are consistent, it means that the person is in a serious input state when filling in invalid input information. At this time, the system will not directly invalidate the web questionnaire. Instead, the system will analyze the situation when the person fills in invalid input information and perform targeted processing.
[0177] First, analyze the proportion of valid content in the web questionnaire content to determine whether invalid input information appears in a large proportion in the web questionnaire content. For invalid input information with a large deviation, it is possible that there is a situation where the person cannot withdraw the questionnaire without submitting it, resulting in a large amount of blank space.
[0178] Step S6031: Based on the proportion of valid content being no greater than the invalidation ratio, obtain the input expression situation of the person when the preset submission button is triggered.
[0179] If the proportion of valid content is no greater than the invalidation ratio, it means that there is less valid input information and invalid input information appears in a large proportion. At this time, the system will obtain the input expression situation of the person when the submission button is triggered.
[0180] Step S60311: When and only when the input expression situation of the person is consistent with the preset panicked input expression, merge all the valid input information in the completed part of the web questionnaire content to form industry questionnaire information and then output it.
[0181] The panicked input expression is a database obtained through training by a neural network model. This database includes images of the facial expression details of people when they are panicked, which will not be elaborated here.
[0182] Compare the input expression situation of the person when the submission button is triggered with the panicked input expression. If the two are inconsistent, then the large amount of invalid input information is deliberately caused by humans, and at this time, the web questionnaire will be directly invalidated.
[0183] If the two are consistent, it means that the person accidentally touched the submission button and submitted the questionnaire before completing the web questionnaire. Therefore, the person will have a panicked expression and there will be a large amount of invalid input information. Since it is not deliberately caused by humans, all the valid input information in the completed part of the web questionnaire content will be retained and all the valid input information will be output.
[0184] Step S6032: Based on the fact that the proportion of valid content is greater than the invalidation ratio, input the invalid input information into a preset sentence correction library for recognition and analysis to obtain corrected input information.
[0185] If the proportion of valid content is greater than the invalidation ratio, it indicates that there is more valid input information. At this time, the appearance of invalid input information may be due to incorrect key presses, resulting in incoherent sentences. For the invalid input information with incoherent sentences, input the invalid input information into the sentence correction library for correction. The sentence correction library is a database set by technicians that can correct incoherent sentences, which will not be elaborated here. The corrected input information refers to the content output after correcting the invalid input information with incoherent sentences.
[0186] Step S60321: Merge the corrected input information and the valid input information to form industry questionnaire information and then output it.
[0187] After correcting the invalid input information, all the content of the questionnaire information is valid input information. At this time, output all the corrected input information and the valid input information.
[0188] Refer to Figure 7 , the method for identifying the content of a paper questionnaire includes the following steps:
[0189] Step S700: Obtain the image information of the paper questionnaire.
[0190] The image information of the paper questionnaire refers to the image obtained by photographing the paper questionnaire with a camera.
[0191] Step S701: According to the image information of the paper questionnaire and the preset questionnaire text specification, frame out all the writing areas in the image information of the paper questionnaire.
[0192] The questionnaire text specification is the standard of the questionnaire content set by technicians, which includes the positions where personnel need to fill in information in the questionnaire content, which will not be elaborated here.
[0193] The writing area refers to the positions where personnel need to fill in information in the questionnaire content.
[0194] By comparing the image information of the paper questionnaire with the questionnaire text specification, all the positions where personnel need to fill in information can be determined from the image information of the paper questionnaire.
[0195] Step S702: Number all the writing areas to obtain the recognition order.
[0196] When the system recognizes the content of the paper questionnaire, it needs to recognize the handwritten Chinese characters in each writing area in turn. The recognition order refers to the order in which the system recognizes the writing areas. The recognition order is determined according to the numbers of each writing area.
[0197] Step S703: extracting the written content of the person in the writing area according to a preset character recognition method in the recognition order.
[0198] The character recognition method refers to a method for recognizing written Chinese characters to extract corresponding Chinese character data, which is common knowledge among those skilled in the art and will not be elaborated here.
[0199] The content written by a person refers to the content filled in by the filling person obtained by identifying the content in the writing area.
[0200] Step S704: If and only if the written content of the person is inconsistent with the preset content format of the writing area, the risk value is invalidated and increased by one, and the accumulated risk value is accumulated in sequence.
[0201] The content format refers to the type of information that needs to be filled in a specific writing area in a paper questionnaire set by the technical staff. If the content filled in by the person filling out the questionnaire in the writing area is inconsistent with the content format, there are errors or random fillings, which will not be elaborated here.
[0202] The invalid risk value is a counting flag, which is used to count when the content written by the personnel is abnormal. The cumulative risk value refers to the final sum of the invalid risk values.
[0203] When the content written by the personnel is inconsistent with the content format of the writing area, it means that there is an abnormality in the content written by the personnel. Every time an abnormality in the content written by the personnel occurs, the risk value of invalidation increases by 1. The higher the cumulative risk value, the higher the risk of the paper questionnaire being invalidated.
[0204] Step S7051: When the cumulative risk value is greater than the preset benchmark risk value, the paper questionnaire is invalidated.
[0205] The benchmark risk value is a reference standard set by technical personnel to measure whether a paper questionnaire is invalidated, and will not be elaborated here.
[0206] When the cumulative risk value is greater than the benchmark risk value, it means that there are many abnormalities in the writing content of people in the writing area, and there is a high probability that people have answered the paper questionnaire randomly. At this time, the paper questionnaire will be directly invalidated.
[0207] Step S7052: When the cumulative risk value is not greater than the benchmark risk value, the written content of the personnel is combined to form the industry questionnaire information and then output.
[0208] If the cumulative risk value is not greater than the baseline risk value, the abnormal writing content in the writing area will be removed, and the writing content in the remaining writing area will be output, and the writing content in the remaining writing area is all normal.
[0209] ReferenceFigure 8 When the writing area is covered with stains, the processing method includes the following steps:
[0210] Step S800: Determine the stain position according to the paper questionnaire image information and the preset stain characteristics.
[0211] The stain position refers to the place on the paper questionnaire where there are stains. By comparing the paper questionnaire image information with the stain characteristics, the stain position can be determined from the paper questionnaire image information. After determining the stain position, the system can perform targeted processing on the stains at the stain position.
[0212] Step S801: Determine the contaminated character content in the personnel's writing content that is contaminated by the stain according to the stain position and the writing area.
[0213] When the stain position is in the writing area, the personnel's writing content in the writing area may be contaminated by the stain characteristics. The contaminated character content refers to the text part in the personnel's writing content that is contaminated by the stain characteristics.
[0214] By comparing the stain position with the writing area, it is possible to determine which positions in the personnel's writing content in the writing area are contaminated by the stain and cannot be directly recognized.
[0215] Step S802: Remove the contaminated character content from the personnel's writing content, and input the remaining personnel's writing content into the preset sentence correction library for correction to obtain the corrected writing text.
[0216] When there are Chinese characters in the personnel's writing content that are contaminated by the stain characteristics and cannot be directly recognized, the contaminated character content is removed from the personnel's writing content, and the remaining personnel's writing content is handed over to the sentence correction library for correction, and there is a possibility of successful correction. The corrected writing text refers to the personnel's writing content corrected by the sentence correction library. There are two situations for the corrected writing text. One situation is that the content is restored completely and the meaning is similar. The other situation is that the meaning of the corrected text has changed and cannot be used. For the above two situations, the system needs to identify them.
[0217] Step S803: Input the corrected writing text and the personnel's writing content in other writing areas into the preset correlation database for correlation matching to obtain a correlation value.
[0218] The correlation database is a database set by technicians for judging the correlation situation of the contaminated personnel's writing content before and after correction, which will not be elaborated here.
[0219] The correlation value refers to the degree of information correlation between the corrected personnel's writing content and the original questionnaire content. The larger the correlation value, the higher the restoration degree of the corrected writing text.
[0220] Since the content in each writing area of the questionnaire has context content associations, the corrected writing text is analyzed and compared with the writing content of other writing areas through the association database to determine the restoration of the corrected writing text.
[0221] Step S8041: When the association value is greater than the preset reference association, the corrected writing text and the writing content of other writing areas are combined to form industry questionnaire information and then output.
[0222] The reference association is a reference value set by technicians to measure the restoration degree of the corrected writing text, which will not be elaborated here.
[0223] When the association value is greater than the reference association, it indicates that the corrected writing text can establish a context relationship with the writing content of other writing areas. The restoration degree of the corrected writing text is relatively high and can be used. At this time, the corrected writing text and the writing content of other writing areas are output.
[0224] Step S8042: When the association value is not greater than the reference association, the corrected writing text is invalidated, and a preset transparent liquid is dropped at the stain position. The handwriting path segment is obtained through the spherical transparent liquid within the preset penetration time.
[0225] If the association value is not greater than the reference association, it means that the restoration degree of the corrected writing text is not high, the meaning is different, and it cannot be used. In this embodiment, by dropping a transparent liquid at the contaminated stain position, when the transparent liquid is at the stain position and is spherical, a part of the stain characteristics at the stain position can move onto the spherical transparent liquid, so that a part of the handwriting path segment of the writing content of the person at the stain position can be presented. The handwriting path segment can be obtained through the transparent liquid, and the spherical liquid can magnify the handwriting path segment. The above method can restore the contaminated writing content of the person.
[0226] The penetration time refers to the time from when the transparent liquid is dropped on the paper questionnaire to when it penetrates and spreads, which will not be elaborated here. The system needs to obtain the handwriting path segment in the transparent liquid within the penetration time.
[0227] The handwriting path segment refers to the handwriting fragments of the writing content of the person that appear in the contaminated writing area after the stain is processed by the transparent liquid, and is obtained by shooting through the transparent liquid with a camera.
[0228] Step S80421: The restored writing text is obtained by inputting the handwriting path segment into a preset character restoration library for restoration.
[0229] The character restoration library is a tool database set for technicians to restore Chinese characters with only partial handwriting fragments, which is common knowledge for those skilled in the art and will not be elaborated here.
[0230] The restored written text refers to the text content obtained by restoring the handwriting path fragments.
[0231] Step S80422: Combine the restored written text and the uncontaminated written content of the person to form industry questionnaire information and then output it.
[0232] After restoring the contaminated written content of the person to obtain the restored written text, combine and output the restored written text and the uncontaminated written content of the person.
[0233] Based on the same inventive concept, an embodiment of the present invention provides a key industry information industry map construction system, including:
[0234] An acquisition module for acquiring the person input behavior image, character information, character input speed, the person input demeanor when filling in the web questionnaire, the person input demeanor when the preset submission button is triggered, the paper questionnaire image information, and the handwriting path fragments.
[0235] A memory for storing a program of a key industry information industry map construction method.
[0236] A processor, and the program in the memory can be loaded and executed by the processor to implement a key industry information industry map construction method.
[0237] Based on the same inventive concept, an embodiment of the present invention provides an intelligent terminal, including a memory and a processor, and a computer program capable of being loaded and executed by the processor to implement a key industry information industry map construction method is stored on the memory.
[0238] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A method for constructing a key industry information industry map, characterized in that: include: Obtain industry questionnaire information using a preset information collection method; Input the industry questionnaire information into the preset keyword library to extract all the keyword information in the industry questionnaire information, and assign a priority weight to each keyword information; Combining single or multiple keyword information to form a keyword combination, and determining the total priority weight of the keyword combination; Input each keyword combination into a preset industry matching database to determine the initial inspection industry information corresponding to the keyword combination; The initial inspection industry information is sorted according to the sum of priority weights to obtain the initial inspection industry information corresponding to the maximum sum of priority weights and defined as key industry information; Determine upstream and downstream position information based on key industry information and preset relevance industry thresholds; Build an industry map based on key industry information, upstream and downstream location information, and preset industry databases.
2. A method for constructing a key industry information industry map according to claim 1, characterized in that: Also includes: Eliminate the key industry information from the initial inspection industry information, and define the remaining initial inspection industry information as auxiliary industry information; Rearrange the auxiliary industry information in descending order according to the sum of priority weights to determine the matching order; According to the matching order and the preset related industry clusters, the auxiliary industry information and the key industry information are tested for similarity to obtain the industry similarity; Based on the industry similarity being greater than a preset similarity threshold, the auxiliary industry upstream and downstream position information is determined according to the auxiliary industry information and the relevant industry threshold; Construct a related industry map based on auxiliary industry information, auxiliary upstream and downstream location information, and industry database; Merge the industry map and the related industry map to get a new industry map.
3. The method for constructing a key industry information industry map according to claim 1, characterized in that: Before combining keyword information, invalid words need to be removed. Invalid words include repeated words and synonyms. The invalid word processing methods include: Classify all keyword information to determine the frequency of occurrence of each keyword information; When there is keyword information with a word frequency greater than 1, the keyword information with a word frequency greater than 1 is defined as a repeated word, and the redundant repeated words are removed and only one is left; After elimination, the information of each pair of keywords is input into a preset synonymous word library for matching to determine whether they are synonymous words; When the two key words are synonymous, the frequency of occurrence of the two key words before elimination is compared; The keyword information with high frequency of occurrence is retained, and the keyword information with low frequency of occurrence is eliminated.
4. The method for constructing a key industry information industry map according to claim 1, characterized in that: Information collection methods include: Obtaining a behavior image input by a person; Determine the input behavior type according to the input behavior image of the person and the preset input tool characteristics, the input behavior type includes web questionnaire input and paper questionnaire input; Based on the webpage questionnaire input, the completed webpage questionnaire content is identified by a preset webpage questionnaire content identification method to obtain the industry questionnaire information and output it; Based on the paper questionnaire input, the completed paper questionnaire content is identified using a preset paper questionnaire content recognition method to obtain industry questionnaire information and output it.
5. A method for constructing a key industry information industry map according to claim 4, characterized in that: Web page questionnaire content recognition methods include: Obtaining character information and character input speed within a preset input recognition area; When the character input speed is greater than the preset normal input speed, the character information is input into a preset character database for analysis to determine whether the character information is recognizable; When the character information is identifiable, the character information is defined as valid input information, otherwise it is defined as invalid input information, and the invalid input information includes blank information, non-text character information and incoherent sentence information; Determine the effective content ratio of the webpage questionnaire content according to the invalid input content and the effective input content; When the proportion of valid content is not greater than the preset invalidation ratio, the webpage questionnaire content will be invalidated; When the proportion of valid content is greater than the invalid proportion, all valid input information in the web page questionnaire content will be combined to form industry questionnaire information and then output.
6. A method for constructing a key industry information industry map according to claim 5, characterized in that: Also includes: When the character input speed is not greater than the normal input speed, determining whether there is invalid input information according to the character information and the character database; If there is invalid input information, obtain the input state of the personnel when filling in the web questionnaire; If and only if the input attitude of the person is consistent with the preset serious input attitude, the effective content ratio of the webpage questionnaire content is determined based on the invalid input content and the valid input content; Based on the fact that the effective content ratio is not greater than the invalid content ratio, the input state of the personnel when the preset submit button is triggered is obtained; If and only if the input state of the person is consistent with the preset panic input state, all valid input information of the completed part of the webpage questionnaire content is combined to form the industry questionnaire information and then output; Based on the fact that the proportion of valid content is greater than the proportion of invalid content, the invalid input information is input into a preset word and sentence correction library for identification and analysis to obtain corrected input information; The corrected input information and the valid input information are combined to form the industry questionnaire information and then output.
7. A method for constructing a key industry information industry map according to claim 4, characterized in that: Paper questionnaire content recognition methods include: Obtain paper questionnaire image information; Select all writing areas in the paper questionnaire image information according to the paper questionnaire image information and the preset questionnaire text specification; Number all writing areas to obtain identification order; Extracting the written content of the person in the writing area according to the recognition order and using a preset character recognition method; If and only if the written content of the person is inconsistent with the preset content format of the writing area, the risk value is invalidated and increased by one, and the cumulative risk value is accumulated in sequence; When the cumulative risk value is greater than the preset benchmark risk value, the paper questionnaire will be invalidated; When the cumulative risk value is not greater than the benchmark risk value, the written content of the personnel will be combined to form the industry questionnaire information and then output.
8. A method for constructing a key industry information industry map according to claim 7, characterized in that: When the written area is covered with stains, the treatment methods include: Determine the stain position according to the paper questionnaire image information and preset stain features; Determine the contaminated character content in the written content of the person that is contaminated by the stain according to the stain position and the writing area; The polluted characters are removed from the written content, and the remaining written content is input into a preset word correction library for correction to obtain a corrected written text; Input the corrected written text and the written content of people in other writing areas into a preset correlation database for correlation matching to obtain a correlation value; When the correlation value is greater than the preset reference correlation, the corrected writing text is combined with the writing content of people in other writing areas to form industry questionnaire information and then output; When the correlation value is not greater than the reference correlation value, the corrected written text is discarded, and a preset transparent liquid is dripped into the stain position, and the handwriting path segment is obtained through the spherical transparent liquid within a preset penetration time; Input the handwriting path fragments into a preset character restoration library for restoration to obtain restored writing text; The restored written text and the uncontaminated written content of personnel are combined to form industry questionnaire information and then output.
9. A key industry information industry map construction system, using a key industry information industry map construction method according to any one of claims 1 to 8, characterized in that: include: An acquisition module is used to acquire the input behavior image of the personnel, character information, character input speed, the input demeanor of the personnel when filling in the web questionnaire, the input demeanor of the personnel when the preset submit button is triggered, the image information of the paper questionnaire, and the handwriting path fragment; A memory for storing a program for a method for constructing a key industry information industry map; The program in the processor and the memory can be loaded and executed by the processor and realize a method for constructing a key industry information industry map.
10. An intelligent terminal, characterized in that: It includes a memory and a processor, and the memory stores a computer program that can be loaded by the processor and execute a method for constructing a key industry information industry map as described in any one of claims 1 to 8.