Intangible asset archive retrieval method and system
By applying information entry and semantic understanding models for intangible asset archives and optimizing search results with user feedback, the problem of inefficient management of traditional intangible asset archives is solved, and efficient and accurate archive retrieval is achieved.
Patent Information
- Application Number
- CN202510141635.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional intangible asset archive management methods are inefficient, which can easily lead to damage or loss of archives, and it is difficult to achieve rapid retrieval and effective utilization, limiting the value of intangible assets.
Provide a method for searching intangible asset archives, which realizes user demand matching and search results optimization through information entry, semantic understanding model, keyword matching and feedback signal adjustment.
The search efficiency and accuracy of intangible asset archives are improved, and the search results are continuously optimized to meet user needs through the training of semantic understanding model and the integration of user feedback.
Smart Images

Figure CN120030147A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an intangible asset archive retrieval method and system. Background Art
[0002] In today's era of rapid development of the knowledge economy, the core competitiveness of enterprises is no longer limited to traditional tangible assets, such as land, equipment and funds, but increasingly relies on the accumulation and use of intangible assets. Intangible assets, as the sum of non-material resources such as corporate innovation achievements, intellectual property rights, brand value and trade secrets, have become a key factor in promoting the sustainable development of enterprises and enhancing their market competitiveness.
[0003] In the related technologies, the traditional intangible asset archive management method mainly relies on manual records and paper document storage, which is not only inefficient but also easily affected by the physical environment, resulting in damage or loss of archives. In addition, manual management makes it difficult to quickly retrieve and effectively utilize a large number of archives, which limits the value of intangible assets and there is room for improvement. Summary of the invention
[0004] In order to improve retrieval efficiency, the present application provides an intangible asset archive retrieval method and system.
[0005] In the first aspect, the present application provides an intangible asset archive retrieval method, which adopts the following technical solution: An intangible asset archive retrieval method, comprising: Enter information into intangible asset archives, obtain archive information, and analyze the archive information to determine archive keywords and corresponding archive overview information; Obtain the user's search statement, and perform semantic understanding on the search statement based on the built-in semantic understanding model to determine the user's needs; Matching the user's needs with the archive keywords and the summary information, determining the archive keywords and the corresponding summary information that match the user's needs in the archive information, and obtaining corresponding first relevance data; Based on the data value of the first relevance data, determine the priority display order of the search results, and output and display the archive keywords and their corresponding summary information according to the priority display order of the search results; Obtaining user feedback signals, and determining the accuracy of search results for each search based on the feedback signals; According to the accuracy of the search results, the user's search statement, the user's needs determined, and the archive keywords matched with accuracy are grouped accordingly to obtain grouping data; Train the semantic understanding model for the user's retrieval statement according to the grouped data to obtain the trained semantic understanding model, and perform semantic understanding on the subsequent user's retrieval statements based on the trained semantic understanding model; Generate a standard retrieval statement based on the accurate retrieval results, and adjust the standard retrieval statement according to the user's retrieval statement to obtain an adjusted retrieval statement with user characteristics and output it together with the corresponding retrieval results.
[0006] Preferably, classify the file information based on the input file information to determine the file category; the file category includes a text category and a picture category; If the file category is the text category, extract keywords from the name data in the file information to obtain name keywords, and use the name keywords as the file keywords of the file; Determine the title data according to the font format of each text character in the file information, and extract keywords from the title data to obtain title keywords; Match the name keywords with the title keywords to determine the correlation degree between the title keywords and the name keywords, and obtain the second correlation degree data; Based on the second correlation degree data, determine the main content in the file information; Summarize the main content to obtain summary information; If the file category is the picture category, perform picture recognition on the file information to determine the picture type; the picture type includes a text scanning type and an image scanning type; If the picture type is the image scanning type, determine the name keywords according to the name data in the file information, and use the name keywords as the file keywords of the file information; If the picture type is the text scanning type, perform image character recognition on the file information to obtain the recognized text data, use the recognized text data as the text data of the file information, and perform text category judgment on the text data of the file information to determine the file keywords and summary information of the file information.
[0007] Preferably, perform character recognition on the image data in the file information based on the image character recognition technology to determine the positive direction of the image, and perform rotation adjustment on the image data according to the positive direction of the image to obtain positive image data; Establish a rectangular coordinate system for the positive image data, and set a detection line at the topmost end of the positive image data; Move the detection line along the ordinate of the rectangular coordinate system from the initial position, and perform image character recognition on the moved area. When text data is recognized, mark the coordinate position of the detection line to obtain marked coordinates; Using the marked coordinates as the starting position of the detection line moving area, the detection line is moved along the ordinate of the rectangular coordinate system and image text recognition is performed on the moved area until the position coordinates of the detection line reach the bottom of the positive image data; Based on the marking coordinates, the recognized text data corresponding to each marking coordinate is counted to obtain line text data; the line text data includes the recognition order of the recognized text and the coordinate data corresponding to the recognized text; Mark the coordinate data of the first recognized character in each line of character data as the first character coordinate, and mark the coordinate data of the last recognized character in each line of character data as the last character coordinate; Perform mathematical statistics on the coordinates of the first character and the last character of each line of text data to determine the special coordinate points in the coordinates of the first character and the special coordinate points in the coordinates of the last character; Match the first character coordinates of each line of text data with the special coordinate point in the first character coordinates, and match the last character coordinates with the special coordinate point in the last character coordinates to determine the line category of the corresponding line of text data; Compare the row categories of two adjacent rows of text data to determine whether the text in the two adjacent rows is continuous, and obtain the adjacent attribute; the adjacent attribute includes continuous and discontinuous; Based on the built-in text template and the adjacent attributes of the corresponding adjacent lines of text data, text generation is performed on the recognized text data to obtain recognized text data.
[0008] Preferably, the valid line interval of the line of text data is determined based on the first character coordinate and the last character coordinate of each line of text data; Counting the valid line intervals, determining the valid line intervals of the entire text in the positive image data, and marking them as valid text intervals; Based on the valid character interval, the position of the special coordinate point in the first character coordinate and the special coordinate point in the last character coordinate are judged to determine the position of the special coordinate point in the valid character interval, and the special coordinate point at the beginning of the valid character interval is marked as the starting point, and the special coordinate point at the end of the valid character interval is marked as the end point; Match the first character coordinates and the last character coordinates in each line of text data with the special coordinate points in the first character coordinates and the special coordinate points in the last character coordinates; If the first character coordinate is a special coordinate point, and the special coordinate point is the starting point, and the last character coordinate is a special coordinate point, and the special coordinate point is the ending point, then the row category of the row of text data is marked as a complete category; If the first character coordinate is a special coordinate point, and the special coordinate point is the starting point, and the last character coordinate is not a special coordinate point, then the row category of the row of text data is marked as the starting category; If the first character coordinate is a special coordinate point, and the special coordinate point is not the starting point, and the last character coordinate is a special coordinate point, and the special coordinate point is the ending point, then the row category of the row of text data is marked as the first category; If the first character coordinate is a special coordinate point, and the special coordinate point is not the starting point, and the last character coordinate is not a special coordinate point, then the row category of the row of text data is marked as the second category; If the first character coordinate is not a special coordinate point, the last character coordinate is a special coordinate point, and the special coordinate point is the end point, then the row category of the row of text data is marked as the end category; If the coordinates of the first character are not a special coordinate point and the coordinates of the last character are not a special coordinate point, the row category of the text data in this row is marked as the center category.
[0009] Preferably, each line of text data is numbered according to the recognition order of the line of text data to obtain numbered data; wherein the line of text data recognized earlier has a smaller numbered data; Based on the order of the number data, the row categories of two adjacent rows are judged. If the row category with the smaller number data is a complete category or the first category, and the row category with the larger number data is a complete category or the starting category, then the adjacent attributes of the two adjacent rows are judged to be continuous; If the row category with the larger number data is the first category, the second category, the middle category, or the last category, then the adjacent attributes of the two adjacent rows are determined to be discontinuous; If the row category with smaller numbered data is the starting category, the second category, the middle category, or the last category, the adjacent attributes of the two adjacent rows are determined to be discontinuous.
[0010] Preferably, receiving a user's feedback signal, and if a user's feedback signal is received, determining the accuracy of the corresponding search result according to the user's feedback signal; If no feedback signal from the user is received, the number of search results is obtained. If the number of search results is 1, the accuracy of the search results is not judged. If the number of search results is greater than 1, the order in which the user views the archives is counted. If the user only views one archive, the viewed information of the viewed archive is counted and judged by big data to determine the accuracy of the archive information; If the user views multiple files, the order in which the user views the files is used to determine whether the user has viewed every file in the search results. If it is determined that the user has viewed every file in the search results, the accuracy of the search results will not be determined. If it is determined that the user has not viewed every file in the search results, the viewed information of the last file viewed will be statistically analyzed and judged based on the user's viewing order to determine the accuracy of the last file viewed.
[0011] Preferably, statistics and judgment are performed on the user's search statements to determine the user's search characteristics; Generate standard search statements corresponding to the search results based on the built-in standard search model and accurate search results; The standard search sentence is adjusted based on the user's search characteristics to obtain an adjusted search sentence with the user's characteristics and output it.
[0012] In the second aspect, the present application provides an intangible asset archive retrieval system, which adopts the following technical solution: An intangible asset archive retrieval system, comprising: an archive data processing module, a user demand judgment module, an archive retrieval matching module, a retrieval result judgment module, a semantic model training module and a retrieval guidance module; The archive data processing module inputs information into the intangible asset archive to obtain archive information, and analyzes the archive information to determine archive keywords and corresponding archive overview information; The user demand judgment module obtains the user's search statement and, based on the built-in semantic understanding model, performs semantic understanding on the search statement to determine the user's demand; The archive search and matching module matches the user's needs with the archive keywords and the summary information, determines the archive keywords and the corresponding summary information in the archive information that match the user's needs, and obtains the corresponding first relevance data; determines the priority display order of the search results based on the data value of the first relevance data, and outputs and displays the archive keywords and the corresponding summary information according to the priority display order of the search results; The search result judgment module obtains the user's feedback signal and determines the accuracy of the search result of each search based on the feedback signal; The semantic model training module groups the user's search statement, the user's needs determined, and the archive keywords matched with accuracy according to the accuracy of the search results to obtain grouping data; performs semantic understanding model training on the user's search statement according to the grouping data to obtain a trained semantic understanding model, and performs semantic understanding on subsequent user's search statements based on the trained semantic understanding model; The search guidance module generates a standard search statement based on accurate search results, and adjusts the standard search statement according to the user's search statement to obtain an adjusted search statement with user characteristics and outputs it together with the corresponding search results.
[0013] In summary, the present application includes at least one of the following beneficial technical effects: By analyzing the archival information of intangible assets, the search keywords of the archival information are determined, so as to facilitate the user's search for archival information. By using the built-in semantic understanding model, the user's search needs are determined, and then the archival information related to the user's search needs is quickly and effectively matched according to the search needs. By collecting user feedback signals, the accuracy of each search result is determined, and then the built-in semantic understanding model is trained according to the accurate search results and the user's search statements, so as to continuously improve the accuracy of the semantic understanding model's understanding of the user's search needs, and then improve the accuracy of the search results matched according to the search needs. At the same time, the standard search statement corresponding to the accurate search result is adjusted according to the user's search statement, so that the user can make adaptive adjustments according to the system's search rules, thereby further improving the user's search accuracy for intangible asset archives and improving the search efficiency; By means of the existing image text recognition technology, the image data is rotated and adjusted, so that the text on the positive image data can be directly recognized. By establishing a rectangular coordinate system for the positive image data and setting detection lines, each line of text on the image is individually divided and recognized, thereby determining the unique coordinates of each text on the image. By performing mathematical statistics on the coordinates of the first and last characters of each line of text, the typing form of the text on the positive image data is determined, and then the adjacent attributes between each line of text are determined. With the help of the adjacent attributes between line text data, the text relationship judgment between lines is simplified, thereby improving the generation efficiency of recognized text data. At the same time, the recognized text is sorted through the built-in text template, which reduces the data confusion of text recognition, thereby improving the accuracy and effectiveness of text sorting, and then the accuracy and effectiveness of the information extracted when the sorted recognized text data is extracted, thereby making the retrieval results more accurate. The coordinates of the first and last characters in each line are used comprehensively to determine the distribution of the characters in the line. By making statistical judgments on the distribution of characters in multiple lines of characters, the valid character interval in the image data for text typing is determined. The character distribution in each line is judged according to the valid character interval to determine the distribution characteristics of the characters in the line, and then the continuity of the characters between lines is clarified, so that the arrangement of the recognized characters is more correct, and the search results are more accurate, thereby improving the user's search effect on archives. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flowchart of the steps of the intangible asset archive retrieval method of this embodiment; Figure 2 This is a module block diagram of the intangible asset archive retrieval system of this embodiment.
[0015] Reference numerals: 1, file data processing module; 2, user requirement judgment module; 3, file retrieval matching module; 4, retrieval result judgment module; 5, semantic model training module; 6, retrieval guidance module. Detailed implementation manners
[0016] The following further elaborates on this application Figure 1-Figure 2 in conjunction with the appended drawings.
[0017] An embodiment of this application discloses a method and system for retrieving intangible asset files.
[0018] Embodiment: As Figure 1 shown, a method for retrieving intangible asset files according to the present invention includes: S100, input information into the intangible asset files to obtain file information, and analyze the file information to determine file keywords and summary information corresponding to the files; wherein, both the file keywords and the summary information are used for retrieval matching, and the summary information also enables users to more quickly and effectively conduct manual screening and confirmation of multiple matching results through display, thereby improving the viewing efficiency of users for target files and further improving the accuracy of retrieval.
[0019] S200, obtain the user's retrieval statement, and based on a built-in semantic understanding model, conduct semantic understanding on the retrieval statement to determine the user's requirements; wherein, the semantic understanding model is a semantic understanding model publicly disclosed in the prior art.
[0020] S300, match the user's requirements with the file keywords and summary information to determine the file keywords in the file information that match the user's requirements and their corresponding summary information, and obtain corresponding first correlation degree data; S400, based on the magnitude of the data value of the first correlation degree data, determine the priority display order of the retrieval results, and output and display the file keywords and their corresponding summary information according to the priority display order of the retrieval results; S500, obtain the user's feedback signal, and determine the accuracy of the retrieval results for each retrieval according to the feedback signal; S600, according to the accuracy of the retrieval results, group the user's retrieval statement, the determined user requirements, and the accurately matched file keywords to obtain grouped data; S700, training the semantic understanding model of the user's search sentence according to the grouping data to obtain the trained semantic understanding model, and semantically understanding the subsequent user's search sentence based on the trained semantic understanding model; wherein the semantic understanding model before training is the semantic understanding model disclosed in the prior art. The trained semantic understanding model is a model for semantically understanding the user's search sentence when searching for intangible asset archives. wherein the training process is a commonly used training method disclosed in the prior art.
[0021] S800, based on the accurate search results, a standard search statement is generated, and the standard search statement is adjusted according to the user's search statement to obtain an adjusted search statement with user characteristics and output it together with the corresponding search results.
[0022] In this embodiment, the archival information of the intangible assets is analyzed to determine the search keywords of the archival information, thereby facilitating the user's retrieval of the archival information. The user's retrieval requirements are determined by utilizing the built-in semantic understanding model, and the archival information related to the user's retrieval requirements is quickly and effectively matched according to the retrieval requirements. The accuracy of each retrieval result is determined by collecting user feedback signals, and the built-in semantic understanding model is trained based on the accurate retrieval results and the user's retrieval statements, thereby continuously improving the accuracy of the semantic understanding model's understanding of the user's retrieval requirements, thereby improving the accuracy of the retrieval results matched according to the retrieval requirements, and at the same time, the standard retrieval statement corresponding to the accurate retrieval result is adjusted according to the user's retrieval statement, thereby enabling the user to make adaptive adjustments based on the system's retrieval rules, thereby further improving the user's retrieval accuracy of intangible asset archives.
[0023] Exemplarily, the intangible asset archives A, B, and C input and stored in the system are analyzed and sorted, and the archive keywords and overview information of the corresponding archives are obtained. For example, the archive keywords of archive A are a1, a2, and a3, and the overview information is Aa. Correspondingly, the archive keywords of archive B are b1, b2, and b3, and the overview information is Bb. The archive keywords of archive C are c1, c2, and c3, and the overview information is Cc. The user demand obtained by the user's search statement is x. By matching x with a1, a2, a3, Aa, b1, b2, b3, Bb, c1, c2, c3, and Cc, the matching result is a match with a1, then it is determined that the archive content that the user wants to view is archive A, and the archive keywords a1, a2, and a3 and the overview information Aa of archive A are displayed for the user to perform manual screening and selection, thereby improving the accuracy of the search results.
[0024] In step S100, information is input into the intangible asset file to obtain the file information, and the file information is analyzed to determine the file keywords and the corresponding file summary information, including the following steps: S110, based on the input file information, classify the file information and determine the file category; the file category includes text category and picture category; the file information of the text category can be patent documents, contracts, etc., and the file information of the picture category can be design drawings, etc.
[0025] S120, if the archive category is a text category, extract keywords from the name data in the archive information to obtain name keywords, and use the name keywords as archive keywords for the archive; S130, determining the title data according to the font format of each text in the archive information, and extracting keywords from the title data to obtain title keywords; S140, matching the name keyword with the title keyword, determining the relevance between the title keyword and the name keyword, and obtaining second relevance data; S150, based on the second correlation data, determining the main content in the archive information; wherein, the higher the second correlation data between the name keyword and the title keyword, the higher the possibility that the content of the corresponding name keyword is the main content.
[0026] S160, summarize the main content to obtain summary information; when summarizing the main content, the existing public technology can be used to summarize the text. At the same time, compared with the existing text summarization technology, by matching the title keywords with the name keywords, the data volume of the text content to be summarized is reduced, thereby improving the summary efficiency.
[0027] S170, if the archive category is a picture category, image recognition is performed on the archive information to determine the picture type; the picture type includes a text scanning type and an image scanning type; when entering archive data, due to the large amount of data to be entered, scanning is often used to digitally convert the organized physical files, so that the original text archives are digitally uploaded in the form of pictures, so the archives of this type of picture category need to be image-text recognized, and the picture archives are converted into text archives for user retrieval. Among them, the text scanning type is an archive of the original text category. When performing digital conversion, the text category archives are converted into picture category archives by using scanning technology, thereby speeding up the digital conversion process and improving the efficiency of information entry. Image scanning types include design drawings, etc.
[0028] S180, if the image type is an image scanning type, determining a name keyword according to the name data in the archive information, and using the name keyword as an archive keyword of the archive information; S190, if the image type is a text scanning type, image text recognition is performed on the archive information to obtain recognized text data, and the recognized text data is used as the text data of the archive information. The text category of the text data of the archive information is judged to determine the archive keywords and summary information of the archive information.
[0029] In this embodiment, by classifying the entered archive information, it is convenient to take more effective technical means to extract the archive keywords and summary information of the archive information, so that the extraction result is more effective and accurate. When the archive category is a text category, the archive keywords and summary information of the archive are clarified by extracting keywords and summarizing the main content of the text content of the archive of the text category. By matching the name keywords and title keywords in the archive information, the most important content in a large amount of text content is determined, thereby reducing the amount of content that needs to be summarized when summarizing the content, thereby improving the summary efficiency, and reducing the interference of other text content on the content that needs to be summarized, thereby improving the accuracy of the summary. When the archive category is a picture category, the archive information of the picture category is re-classified to further clarify the more specific category of the archive information of the picture category. By performing image text recognition on the picture category archive of the text scanning type, the accuracy of the information extraction of the archive information is further guaranteed.
[0030] By dividing the entered archival information into categories twice, the archival category of the archival information is clarified, and corresponding measures are taken according to different archival categories to enable accurate and effective information extraction of the archival information. By matching the name keywords of the text category with the title keywords, the main text content in a large amount of text data is determined, thereby reducing the text content that needs to be summarized, improving the summary efficiency and accuracy. By performing image text recognition on the archival information of the text scanning type, the information extraction of the archival information is more comprehensive, so that when matching user needs according to the archival keywords and the summary information, the matching results are more in line with the user's needs, thereby improving the accuracy of the retrieval.
[0031] In step S190, image text recognition is performed on the archive information to obtain recognized text data, including the following steps: S191, based on the image text recognition technology, the image data in the archive information is subjected to text recognition, the positive direction of the image is determined, and the image data is rotated and adjusted according to the positive direction of the image to obtain positive image data; wherein, the positive direction of the image is the direction of the image when the text can be directly recognized after the image is corrected for text.
[0032] S192, establish a rectangular coordinate system for the positive image data, and set up a detection line at the top of the positive image data; for example, the positive image data is a rectangle, for example, the four vertices are A, B, C, and D, and point A is used as the origin of the rectangular coordinate system. The coordinates of the four vertices are A (0, 0), B (10, 0), C (0, 20), and D (10, 20), respectively. Then, a detection line is set up at the top of the positive image data, and the initial position of the detection line is y=20, where y is the ordinate of the rectangular coordinate system.
[0033] S193, move the detection line from the initial position along the ordinate of the rectangular coordinate system, and perform image text recognition on the moved area. When text data is recognized, mark the coordinate position of the detection line to obtain the marked coordinates; by gradually changing the position of the detection line, perform text recognition on the area between the two positions. For example, the initial position of the detection line is y=20, and the value of each movement is 1. After the first movement, the area that needs to be recognized is (0,19)~(10,20).
[0034] S194, using the marked coordinates as the starting position of the detection line moving area, the detection line is moved along the ordinate of the rectangular coordinate system and image text recognition is performed on the moving area until the position coordinate of the detection line reaches the bottom of the positive image data; when the text recognition area (0,19) to (10,20) recognizes text, the second text recognition area is (0,18) to (10,19); if the text recognition area is (0,18) to (10,19) and no text is recognized, the third text recognition area is (0,17) to (10,19), and so on, until the coordinate position of the detection line is y=0.
[0035] S195, based on the mark coordinates, statistics are collected on the recognized text data corresponding to each mark coordinate to obtain line text data; the line text data includes the recognition order of the recognized text and the coordinate data corresponding to the recognized text; wherein the line text data is a general term for the text data of each line.
[0036] S196, marking the coordinate data of the first recognized character in each line of character data as the first character coordinate, and marking the coordinate data of the last recognized character in each line of character data as the last character coordinate; S197, performing mathematical and statistical judgment on the coordinates of the first character and the last character of each line of text data, and determining the special coordinate points in the first character coordinates and the special coordinate points in the last character coordinates; wherein the special coordinate points can be the starting position of a line, i.e., the top line of writing, the ending position of a line, or the first line position of a paragraph with "the first line indented by two characters".
[0037] S198, matching the first character coordinates of each line of text data with a special coordinate point in the first character coordinates, and matching the last character coordinates with a special coordinate point in the last character coordinates, to determine the line category of the corresponding line of text data; S199, compare the row categories of two adjacent lines of text data, determine whether the text of the two adjacent lines is continuous, and obtain the adjacent attribute; the adjacent attribute includes continuous and discontinuous; number each line of text data according to the recognition order of the line of text data to obtain the number data; wherein, the earlier the line of text data is recognized, the smaller the number data corresponding to; based on the order of the number data, judge the row categories of the two adjacent lines, if the row category with the smaller number data is the complete category or the first category, and the row category with the larger number data is the complete category or the starting category, then the adjacent attribute of the two adjacent lines is determined to be continuous; if the row category with the larger number data is the first category or the second category or the middle category or the last category, then the adjacent attribute of the two adjacent lines is determined to be discontinuous; if the row category with the smaller number data is the starting category or the second category or the middle category or the last category, then the adjacent attribute of the two adjacent lines is determined to be discontinuous. With the help of the adjacent attribute between the line of text data, the text relationship judgment between lines is simplified, thereby improving the generation efficiency of the recognized text data.
[0038] S1910, based on the built-in text template and the adjacent attributes of the corresponding adjacent lines of text data, text generation is performed on the recognized text data to obtain recognized text data.
[0039] In this embodiment, when performing image text recognition on archival information, the positive direction of the text on the image data in the archival information is determined by utilizing existing image text recognition technology, and the image data is rotated and adjusted so that the text on the positive image data can be directly recognized. By establishing a rectangular coordinate system for the positive image data and setting detection lines, each line of text on the image is individually divided and recognized, thereby determining the unique coordinates of each text on the image. By performing mathematical statistics on the coordinates of the first and last characters of each line of text, the typing form of the text on the positive image data is determined, and then the adjacent attributes between each line of text are determined. Then, the recognized text is sorted through the built-in text template, thereby reducing data confusion in text recognition, thereby improving the accuracy and effectiveness of text sorting, and thereby making the accuracy and effectiveness of the information extracted when extracting information from the sorted recognized text data, thereby making the retrieval results more accurate.
[0040] In step S198, the first character coordinates of each line of text data are matched with the special coordinate points in the first character coordinates, and the last character coordinates are matched with the special coordinate points in the last character coordinates to determine the line category of the corresponding line of text data, including the following steps: S1981, determining a valid line interval of the line of text data based on the coordinates of the first character and the last character of each line of text data; S1982, count the valid line intervals, determine the valid line intervals of the entire text in the positive image data, and mark them as valid text intervals; for example, the coordinates of the positive image data in the rectangular coordinate system are (0,0) to (10,20), and the actual text input area is (1,1) to (9,19), so the valid text interval is 1 to 9.
[0041] S1983, based on the valid character interval, determine the position of the special coordinate point in the first character coordinate and the special coordinate point in the last character coordinate, determine the position of the special coordinate point in the valid character interval, mark the special coordinate point at the beginning of the valid character interval as the starting point, and mark the special coordinate point at the end of the valid character interval as the ending point; S1984, matching the first character coordinates and the last character coordinates in each line of character data with the special coordinate points in the first character coordinates and the special coordinate points in the last character coordinates; S1985, if the first character coordinate is a special coordinate point, and the special coordinate point is the starting point, and the last character coordinate is a special coordinate point, and the special coordinate point is the ending point, then the row category of the text data in this row is marked as a complete category; for example, the effective text length is divided into three sections, corresponding to 0, 1, 2, and 3 respectively, where 0~3 is a complete category, 0~x is a starting category (x is less than 3), 1~3 is the first category, 1~x is the second category; 1~2 is a center category; x~3 is an ending category.
[0042] S1986, if the first character coordinate is a special coordinate point, and the special coordinate point is the starting point, and the last character coordinate is not a special coordinate point, then mark the row category of the row of character data as the starting category; S1987, if the first character coordinate is a special coordinate point, and the special coordinate point is not the starting point, and the last character coordinate is a special coordinate point, and the special coordinate point is the ending point, then the row category of the text data in this row is marked as the first category; for example, a text with a two-character indent in the first line occupies the entire line.
[0043] S1988, if the first character coordinate is a special coordinate point, and the special coordinate point is not the starting point, and the last character coordinate is not a special coordinate point, then the line category of the text data in this line is marked as the second category; for example, a paragraph with a two-character indent on the first line does not occupy the entire line.
[0044] S1989, if the coordinates of the first character are not a special coordinate point, the coordinates of the last character are a special coordinate point, and the special coordinate point is the end point, then the row category of the row of text data is marked as the end category; for example, left-aligned, etc.
[0045] S19810: If the coordinates of the first character are not a special coordinate point and the coordinates of the last character are not a special coordinate point, the row category of the text data in this row is marked as a centered category.
[0046] In this embodiment, the distribution of the text in each line is determined by the coordinates of the first and last characters in the line. The text distribution of multiple lines of text is statistically determined to determine the valid text interval in the image data for text typing. The text distribution in each line is judged based on the valid text interval to determine the distribution characteristics of the text in the line, and then the continuity of the text between lines is clarified, so that the arrangement of the recognized text is more correct, and the search results are more accurate, thereby improving the user's search effect on archives.
[0047] Exemplarily, the area of the positive image data in the rectangular coordinate system is (0, 0) to (10, 20), while the actual text input area is (1, 1) to (9, 19). Therefore, the effective text interval is from 1 to 9. Due to the continuity of the text, the coordinates of the first recognized text in each line of the positive image data are mostly the starting positions of the effective text interval. Thus, by counting the coordinates of the first text in each line of text data, the starting point of the effective text interval of the positive image data is determined. Similarly, the coordinates of the last recognized text in each line of the positive image data are mostly the ending positions of the effective text interval. Therefore, by counting the coordinates of the last text in each line of text data, the ending point of the effective text interval of the positive image data is determined.
[0048] When the starting point is the coordinate of the first text in a certain line and the ending point is the coordinate of the last text, it indicates that the text distribution in this line is in the text format from start to end, that is, the complete category. Similarly, when the starting point is the coordinate of the first text in a certain line, and the ending point is not the ending point, it indicates that the text distribution in this line is from the start to the end of a certain area. Furthermore, it indicates that there must be a text segmentation phenomenon between this line and the next line. Therefore, it is marked as the starting category. When the starting point of a certain line is not the starting point but is a special coordinate point, such as the position of the first line of a paragraph with the first line indented by two characters, it indicates that there is a special text format in this line and it is necessary to indent the first line, etc. Furthermore, it indicates that there must be a text segmentation phenomenon between this line and the previous line. And if in this case, the ending point is the ending point, it indicates that the effective text interval is fully input in the special text format. Furthermore, it indicates that there may be a continuous phenomenon between this line and the next line. Therefore, it is marked as the first category; on the contrary, if in this case, the ending point is not the ending point, it indicates that the effective text interval is not fully input in the special text format. Furthermore, it indicates that there must be a text segmentation phenomenon between this line and the next line. Therefore, it is marked as the second category. When the starting point of a certain line is not a special coordinate point and the ending point is the ending point, it indicates that there is a special text format in this line, such as text signature, right alignment, etc. Therefore, it is marked as the ending category. When both the starting point and the ending point of a certain line are not special coordinate points, it indicates that there is a special text format in the text of this line, such as centered alignment. Therefore, it is marked as the centered category.
[0049] In step S500, obtain the feedback signal of the user and determine the accuracy of the retrieval result of each retrieval according to the feedback signal, including the following steps: S510, receiving a user's feedback signal. If a user's feedback signal is received, determining the accuracy of the corresponding search result according to the user's feedback signal; if the user's feedback signal indicates that the search result is accurate, then determining that the corresponding search result is accurate. For example, if the user's feedback signal indicates that search result A is accurate, then determining that search result A is accurate. If the user's feedback signal indicates that search result A is inaccurate, then determining that search result A is inaccurate.
[0050] S520, if no feedback signal from the user is received, the number of search results is obtained. If the number of search results is 1, the accuracy of the search results is not judged. When the user does not provide feedback on the search results, it is necessary to analyze and judge the user's viewing situation to determine the user's satisfaction with the search results, and then determine the accuracy of the corresponding search results.
[0051] S530, if the number of search results is greater than 1, then the order in which the user views the archives is counted; if the user views only one archive, then the viewed information of the viewed archive is counted and judged by big data to determine the accuracy of the archive information; S540, if the user views multiple files, determining whether the user has viewed each file in the search results according to the order in which the user views the files, and if it is determined that the user has viewed each file in the search results, no accuracy judgment is performed on the search results; S550, if it is determined that the user has not viewed every file in the search result, statistics and big data judgment are performed on the viewed information of the last viewed file according to the viewing order of the user to determine the accuracy of the last viewed file information.
[0052] In this embodiment, the accuracy of the search results is judged in the simplest and most effective way by collecting user feedback information. If the user does not provide feedback on the search results, the number of search results is judged to determine whether the search results need to be judged for accuracy, so as to simplify the judgment process. When the number of search results is greater than 1, the user's viewing order is judged to clarify the search results that need to be judged for accuracy, thereby reducing the amount of data to be judged and improving the judgment efficiency.
[0053] Exemplarily, if a user's feedback result is received, for example, if the user's feedback signal indicates that the search result A is accurate, then the search result A is determined to be accurate; if the user's feedback signal indicates that the search result A is inaccurate, then the search result A is determined to be inaccurate.
[0054] If no feedback is received from the user, the quantitative value of the search results is judged. For example, if the search result only contains file A, then since there is no reference and it is unique, the next time the user performs the same search, the search result will be unique, so there is no need to judge the accuracy. Furthermore, since there is no reference, it is impossible to determine the user's opinion on the search result, so no accuracy judgment is performed.
[0055] If the search results include files A, B, and C, and the user only views file A, then by counting the number of times the user views files A, B, and C and making a big data judgment, for example, if there are multiple simultaneous occurrences of files A, B, and C in a large number of search results of the user, and the user views most of the three files as file A, then it is determined that file A better meets the user's needs, and thus file A is more accurate; However, if the user views multiple files among files A, B, and C, for example, A, B, and C, then since there is no distinctiveness, it is impossible to determine the user's demand tendency for file viewing by the number of times A, B, and C are viewed, so it is impossible to make an accuracy judgment.
[0056] If the user views files A and B in sequence but does not view file C, it means that the user's needs were not met when viewing file A, but were met when viewing file B. Therefore, the accuracy of file B is greater than that of file A. Therefore, by performing big data analysis on the viewing of file B, the accuracy of file B can be further determined, making the accuracy judgment of file B more accurate.
[0057] In step S800, based on the accurate search results, a standard search statement is generated, and the standard search statement is adjusted according to the user's search statement to obtain an adjusted search statement with user characteristics and output it together with the corresponding search results, including the following steps: S810, performing statistics and judgment on the user's search statements to determine the user's search characteristics; S820, generating a standard search statement corresponding to the search result based on the built-in standard search model and the accurate search result; S830, adjusting the standard search sentence based on the user's search characteristics, obtaining an adjusted search sentence with the user's characteristics, and outputting the adjusted search sentence.
[0058] In this embodiment, the user's search statements are sorted and judged to determine the user's search characteristics when performing demand search. Then, when the user performs the next demand search, a corresponding standard search statement is generated based on the search results, and the standard search statement is adjusted and output according to the user's search characteristics. When the user determines that the search results are accurate, the user can adjust the search statement independently by learning how to adjust the search statement, so that the system can understand the user's search statement more accurately and effectively, which not only corrects the user's search method, but also improves the accuracy of the semantic understanding model's understanding of the user's search statement, thereby making the search results more accurate.
[0059] Based on the description of the above intangible asset archive retrieval method embodiment, the embodiment of the present invention also discloses an intangible asset archive retrieval system: like Figure 2 As shown, an intangible asset archive retrieval system, by applying the above-mentioned intangible asset archive retrieval method, includes: an archive data processing module 1, a user demand judgment module 2, an archive retrieval matching module 3, a retrieval result judgment module 4, a semantic model training module 5 and a retrieval guidance module 6; The archive data processing module 1 inputs information into the intangible asset archive, obtains archive information, and analyzes the archive information to determine archive keywords and corresponding archive overview information; User demand judgment module 2 obtains the user's search sentence and, based on the built-in semantic understanding model, performs semantic understanding on the search sentence to determine the user's demand; The archive search matching module 3 matches the user demand with the archive keywords and the summary information, determines the archive keywords and the corresponding summary information in the archive information that match the user demand, and obtains the corresponding first relevance data; determines the priority display order of the search results based on the data value of the first relevance data, and outputs and displays the archive keywords and the corresponding summary information according to the priority display order of the search results; The search result judgment module 4 obtains the user's feedback signal and determines the accuracy of the search result of each search according to the feedback signal; The semantic model training module 5 groups the user's search statement, the user's needs determined, and the archive keywords matched with accuracy according to the accuracy of the search results to obtain grouping data; performs semantic understanding model training on the user's search statement according to the grouping data to obtain a trained semantic understanding model, and performs semantic understanding on subsequent user's search statements based on the trained semantic understanding model; The search guidance module 6 generates a standard search statement based on the accurate search results, and adjusts the standard search statement according to the user's search statement to obtain an adjusted search statement with user characteristics and outputs it together with the corresponding search results.
[0060] Compared with the existing intangible asset archive retrieval method and system, the present invention improves the retrieval efficiency.
[0061] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.
Claims
1. An intangible asset archive retrieval method, characterized in that: include: Enter information into intangible asset archives, obtain archive information, and analyze the archive information to determine archive keywords and corresponding archive overview information; Obtain the user's search statement, and perform semantic understanding on the search statement based on the built-in semantic understanding model to determine the user's needs; Matching the user's needs with the archive keywords and the summary information, determining the archive keywords and the corresponding summary information that match the user's needs in the archive information, and obtaining corresponding first relevance data; Based on the data value of the first relevance data, determine the priority display order of the search results, and output and display the archive keywords and their corresponding summary information according to the priority display order of the search results; Obtaining user feedback signals, and determining the accuracy of search results for each search based on the feedback signals; According to the accuracy of the search results, the user's search statement, the user's needs determined, and the archive keywords matched with accuracy are grouped accordingly to obtain grouping data; Performing semantic understanding model training on the user's search sentences according to the grouping data to obtain a trained semantic understanding model, and performing semantic understanding on subsequent user's search sentences based on the trained semantic understanding model; Based on accurate search results, a standard search statement is generated, and the standard search statement is adjusted according to the user's search statement to obtain an adjusted search statement with user characteristics and output it together with the corresponding search results.
2. The intangible asset archive retrieval method according to claim 1, characterized in that: The intangible asset file is input to obtain file information, and the file information is analyzed to determine file keywords and corresponding file overview information, specifically: Based on the input archive information, classify the archive information to determine the archive category; the archive category includes text category and picture category; If the archive category is a text category, keyword extraction is performed on the name data in the archive information to obtain name keywords, and the name keywords are used as archive keywords of the archive; According to the font format of each text in the archive information, the title data is determined, and keywords are extracted from the title data to obtain the title keywords; Matching the name keyword with the title keyword, determining the relevance between the title keyword and the name keyword, and obtaining second relevance data; Determining the main content of the archive information based on the second relevance data; Summarize the main content to obtain summary information; If the archive category is a picture category, image recognition is performed on the archive information to determine the picture type; the picture type includes a text scanning type and an image scanning type; If the image type is an image scanning type, determine the name keyword based on the name data in the archive information, and use the name keyword as the archive keyword of the archive information; If the image type is a text scanning type, image text recognition is performed on the archive information to obtain recognized text data, and the recognized text data is used as the text data of the archive information. The text category of the text data of the archive information is judged to determine the archive keywords and overview information of the archive information.
3. The intangible asset archive retrieval method according to claim 2, characterized in that: The image and text recognition of the archive information is performed to obtain the recognized text data, specifically: Based on the image and text recognition technology, the image data in the archive information is subjected to text recognition, the positive direction of the image is determined, and the image data is rotated and adjusted according to the positive direction of the image to obtain positive image data; Establish a rectangular coordinate system for the positive image data, and set up a detection line at the top of the positive image data; The detection line is moved from the initial position along the ordinate of the rectangular coordinate system, and image text recognition is performed on the moved area. When text data is recognized, the coordinate position of the detection line is marked to obtain the marked coordinates; Using the marked coordinates as the starting position of the detection line moving area, the detection line is moved along the ordinate of the rectangular coordinate system and image text recognition is performed on the moved area until the position coordinates of the detection line reach the bottom of the positive image data; Based on the marking coordinates, the recognized text data corresponding to each marking coordinate is counted to obtain line text data; the line text data includes the recognition order of the recognized text and the coordinate data corresponding to the recognized text; Mark the coordinate data of the first recognized character in each line of character data as the first character coordinate, and mark the coordinate data of the last recognized character in each line of character data as the last character coordinate; Perform mathematical statistics on the coordinates of the first character and the last character of each line of text data to determine the special coordinate points in the coordinates of the first character and the special coordinate points in the coordinates of the last character; Match the first character coordinates of each line of text data with the special coordinate point in the first character coordinates, and match the last character coordinates with the special coordinate point in the last character coordinates to determine the line category of the corresponding line of text data; Compare the row categories of two adjacent rows of text data to determine whether the text in the two adjacent rows is continuous, and obtain the adjacent attribute; the adjacent attribute includes continuous and discontinuous; Based on the built-in text template and the adjacent attributes of the corresponding adjacent lines of text data, text generation is performed on the recognized text data to obtain recognized text data.
4. The intangible asset archive retrieval method according to claim 3, characterized in that: The first character coordinates of each line of text data are matched with the special coordinate points in the first character coordinates, and the last character coordinates are matched with the special coordinate points in the last character coordinates to determine the line category of the corresponding line of text data, specifically: Determine the valid line interval of the line text data based on the first character coordinate and the last character coordinate of each line text data; Counting the valid line intervals, determining the valid line intervals of the entire text in the positive image data, and marking them as valid text intervals; Based on the valid character interval, the position of the special coordinate point in the first character coordinate and the special coordinate point in the last character coordinate are judged to determine the position of the special coordinate point in the valid character interval, and the special coordinate point at the beginning of the valid character interval is marked as the starting point, and the special coordinate point at the end of the valid character interval is marked as the end point; Match the first character coordinates and the last character coordinates in each line of text data with the special coordinate points in the first character coordinates and the special coordinate points in the last character coordinates; If the first character coordinate is a special coordinate point, and the special coordinate point is the starting point, and the last character coordinate is a special coordinate point, and the special coordinate point is the ending point, then the row category of the row of text data is marked as a complete category; If the first character coordinate is a special coordinate point, and the special coordinate point is the starting point, and the last character coordinate is not a special coordinate point, then the row category of the row of text data is marked as the starting category; If the first character coordinate is a special coordinate point, and the special coordinate point is not the starting point, and the last character coordinate is a special coordinate point, and the special coordinate point is the ending point, then the row category of the row of text data is marked as the first category; If the first character coordinate is a special coordinate point, and the special coordinate point is not the starting point, and the last character coordinate is not a special coordinate point, then the row category of the row of text data is marked as the second category; If the first character coordinate is not a special coordinate point, the last character coordinate is a special coordinate point, and the special coordinate point is the end point, then the row category of the row of text data is marked as the end category; If the coordinates of the first character are not a special coordinate point and the coordinates of the last character are not a special coordinate point, the row category of the text data in this row is marked as the center category.
5. The intangible asset archive retrieval method according to claim 4, characterized in that: The row categories of two adjacent rows of text data are compared to determine whether the text of the two adjacent rows is continuous, and the adjacent attributes are obtained, which are specifically: Numbering each line of text data according to the recognition order of the line of text data to obtain numbered data; wherein the line of text data recognized earlier has a smaller numbered data; Based on the order of the number data, the row categories of two adjacent rows are judged. If the row category with the smaller number data is a complete category or the first category, and the row category with the larger number data is a complete category or the starting category, then the adjacent attributes of the two adjacent rows are judged to be continuous; If the row category with the larger number data is the first category, the second category, the middle category, or the last category, then the adjacent attributes of the two adjacent rows are determined to be discontinuous; If the row category with smaller numbered data is the starting category, the second category, the middle category, or the last category, the adjacent attributes of the two adjacent rows are determined to be discontinuous.
6. The intangible asset archive retrieval method according to claim 1, characterized in that: The step of obtaining a user's feedback signal and determining the accuracy of the search result of each search according to the feedback signal is specifically as follows: Receiving a user's feedback signal, and if a user's feedback signal is received, determining the accuracy of the corresponding search result according to the user's feedback signal; If no feedback signal from the user is received, the number of search results is obtained. If the number of search results is 1, the accuracy of the search results is not judged. If the number of search results is greater than 1, the order in which the user views the archives is counted. If the user only views one archive, the viewed information of the viewed archive is counted and judged by big data to determine the accuracy of the archive information; If the user views multiple files, the order in which the user views the files is used to determine whether the user has viewed every file in the search results. If it is determined that the user has viewed every file in the search results, the accuracy of the search results will not be determined. If it is determined that the user has not viewed every file in the search results, the viewed information of the last file viewed will be statistically analyzed and judged based on the user's viewing order to determine the accuracy of the last file viewed.
7. The intangible asset archive retrieval method according to claim 1, characterized in that: The method generates a standard search statement based on the accurate search results, and adjusts the standard search statement according to the user's search statement to obtain an adjusted search statement with user characteristics and outputs it together with the corresponding search results, specifically: Count and judge the user's search statements to determine the user's search characteristics; Generate standard search statements corresponding to the search results based on the built-in standard search model and accurate search results; The standard search sentence is adjusted based on the user's search characteristics to obtain an adjusted search sentence with the user's characteristics and output it.
8. An intangible asset archive retrieval system, characterized in that: The system is used to implement an intangible asset archive retrieval method as described in any one of claims 1 to 7: comprising: an archive data processing module, a user demand judgment module, an archive retrieval matching module, a retrieval result judgment module, a semantic model training module and a retrieval guidance module; The archive data processing module inputs information into the intangible asset archive to obtain archive information, and analyzes the archive information to determine archive keywords and corresponding archive overview information; The user demand judgment module obtains the user's search sentence and, based on the built-in semantic understanding model, performs semantic understanding on the search sentence to determine the user's demand; The archive search and matching module matches the user's needs with the archive keywords and the summary information, determines the archive keywords and the corresponding summary information in the archive information that match the user's needs, and obtains the corresponding first relevance data; determines the priority display order of the search results based on the data value of the first relevance data, and outputs and displays the archive keywords and the corresponding summary information according to the priority display order of the search results; The search result judgment module obtains the user's feedback signal and determines the accuracy of the search result of each search according to the feedback signal; The semantic model training module groups the user's search statement, the user's needs determined, and the archive keywords matched with accuracy according to the accuracy of the search results to obtain grouping data; performs semantic understanding model training on the user's search statement according to the grouping data to obtain a trained semantic understanding model, and performs semantic understanding on subsequent user's search statements based on the trained semantic understanding model; The search guidance module generates a standard search statement based on accurate search results, and adjusts the standard search statement according to the user's search statement to obtain an adjusted search statement with user characteristics and outputs it together with the corresponding search results.