Method and device for extracting structured information from text
By calculating the attribute value degree of text and the user behavior value degree, selecting regular expressions with appropriate granularity, extracting structured information from unstructured text and storing it in the corresponding system, solving the problem of storage difficulties in the prior art and realizing the classified storage of text information.
Patent Information
- Application Number
- CN202510829616.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The prior art is difficult to realize the classified storage of structured information extracted from unstructured text, resulting in storage difficulties.
By calculating the comprehensive value score based on the attribute value degree of the text and the user behavior value degree, select the corresponding granularity regular expression, extract structured information from the text, and store it in the corresponding structured storage system.
Effective classification storage of structured information in unstructured text is realized, and the storage difficulties in the prior art are solved.
Smart Images

Figure CN120336530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and apparatus for extracting structured information from text. Background Art
[0002] During the daily business handling process of enterprise employees, business data is usually recorded in the form of unstructured text. However, when it is necessary to store these business data, the storage system is usually a structured storage system. Therefore, in this storage process, it is crucial to extract business data from the unstructured text and convert it into structured information.
[0003] Currently, neural networks or regular expressions are usually used to extract business data from unstructured text and convert it into structured information. For example, Patent 202111279596.X discloses a method for extracting text structured information based on a neural network and its related devices, which provides a method for extracting structured information from text through a neural network. Patent CN202110701621.2 discloses a method, device, and storage medium for accurately extracting complex web page structured information, which discloses a method for extracting structured information through regular expressions.
[0004] However, considering that structured storage systems usually need to classify and store structured information, the above various methods for extracting structured information can only directly extract structured information from text, resulting in difficulties in subsequent storage. For example, it is easy to have difficulties in determining the structured information to be stored in the structured storage system. For example, in practical applications, structured information is usually stored in multiple structured storage systems respectively. Currently, the extraction method directly uses neural networks and regular expressions to extract structured information from text. However, even in the same text, there are usually a large number of data blocks, and the structured information extracted from these data blocks may need to be stored in different structured storage systems. Obviously, the current extraction method is difficult to achieve this classification storage. Summary of the Invention
[0005] The present invention provides a method and apparatus for extracting structured information from text, aiming to solve the problem that it is difficult to achieve classification storage when extracting structured information from text in the prior art.
[0006] On the one hand, the present invention provides a method for extracting structured information from text, including: Determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted; For the target data block in the text to be extracted, determine the user behavior value degree of the target data block according to the user operation record of the target data block; wherein, the target data block is specifically any data block in the text to be extracted; Determine the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block; According to the comprehensive value score of the target data block, select a regular expression corresponding to the extraction granularity, and extract structured information from the target data block; Store the extracted structured information in a structured storage system corresponding to the comprehensive value score.
[0007] Preferably, the method further includes presetting regular expressions with coarse extraction granularity, regular expressions with fine extraction granularity, and a preset threshold; and, According to the comprehensive value score of the target data block, selecting a regular expression corresponding to the extraction granularity and extracting structured information from the target data block specifically includes: Judge whether the comprehensive value score of the target data block is greater than the preset threshold; When the comprehensive value score is greater than the preset threshold, select a regular expression with fine extraction granularity and extract structured information from the target data block; or, When the comprehensive value score is less than or equal to the preset threshold, select a regular expression with coarse extraction granularity and extract structured information from the target data block.
[0008] Preferably, determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted specifically includes: Obtain the internal importance level of the text to be extracted, and determine an importance score according to the internal importance level; wherein, the importance score is used to represent the importance of the text to be extracted within the enterprise; Obtain the network popularity of the content classification label, and determine a popularity score according to the network popularity; wherein, the popularity score is used to represent the popularity of the content classification label of the text to be extracted in online media; Determine the attribute value degree of the text to be extracted according to the importance score and the popularity score.
[0009] Preferably, determining the attribute value degree of the text to be extracted according to the importance score and the popularity score specifically includes calculating the attribute value degree through the following formula: ; Wherein: the P attribute is the attribute value degree; I is the importance score; H is the popularity score; both δ and γ are adjustment coefficients less than 1, and γ is greater than δ.
[0010] Preferably, according to the user operation records of the target data block, determine the user behavior value degree of the target data block, specifically including: Determine the user behavior value degree of the target data block according to the modification frequency of the target data block, the time interval since the last modification of the target data block, and the average viewing duration of the target data block.
[0011] Preferably, according to the modification frequency of the target data block, the time interval since the last modification of the target data block, and the average viewing duration of the target data block, determine the user behavior value degree of the target data block, specifically including: Determine the data activity score according to the modification frequency, determine the user attention score according to the average viewing duration, and determine the data freshness of the target data block according to the time interval; Perform a weighted sum of the data activity score and the user attention score, and multiply the result of the weighted sum by the data freshness to obtain the user behavior value degree.
[0012] Preferably, determine the data activity score according to the modification frequency, specifically including: Obtain the maximum value and the minimum value among the modification frequencies of each data block in the text to be extracted; Calculate the data activity score by passing the modification frequency of the target data block, the maximum value, and the minimum value through a normalization calculation formula.
[0013] Preferably, determine the data freshness of the target data block according to the time interval, specifically including: Judge whether the time interval is greater than the data validity period set for the target data block; If so, determine the data freshness as 0; or, If not, calculate the data freshness through the following formula; ; Wherein, T is the data freshness; t is the time interval; t 有效期 is the data validity period set for the target data block.
[0014] Preferably, determine the user attention score according to the average viewing duration, specifically including: Obtain the average viewing duration of the text to be extracted; Through the formula to determine the user importance score, where D is the user importance score; min is the operator for taking the minimum value; d is the average access duration of the target data block; d total is the average access duration of the text to be extracted.
[0015] On the other hand, the present invention provides an apparatus for extracting structured information from text, including: A first determination unit, configured to determine the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted; A second determination unit, configured to determine the user behavior value degree of the target data block in the text to be extracted according to the user operation record of the target data block; wherein the target data block is specifically any data block in the text to be extracted; A third determination unit, configured to determine the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block; An extraction unit, configured to select a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block, and extract structured information from the target data block; A storage unit, configured to store the extracted structured information into a structured storage system corresponding to the comprehensive value score.
[0016] The method for extracting structured information from text provided by the embodiments of the present application includes determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted, determining the user behavior value degree of the target data block in the text to be extracted according to the user operation record of the target data block, determining the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block, then selecting a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block, extracting structured information from the target data block, and then storing the extracted structured information into a structured storage system corresponding to the comprehensive value score. This method calculates the comprehensive value score of the target data block in the text to be extracted to select a regular expression corresponding to the extraction granularity, extracts structured information from the target data block, and stores the structured information into the structured storage system corresponding to the comprehensive value score. Therefore, even if there are multiple data blocks in the text to be extracted, structured information can be extracted from each data block through this method and stored into the corresponding structured storage system, thus solving the problems in the prior art. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flow chart of a method for extracting structured information from text provided by the present invention; Figure 2 It is a structural block diagram of a device for extracting structured information from text provided by the present invention; Figure 3 It is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners
[0019] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0020] As mentioned above, the current extraction methods using neural networks and regular expressions directly extract text structured information from text. However, considering that structured storage systems usually require classified storage of structured information, it is easy to have a situation where it is difficult to determine the structured storage system to which the extracted structured information needs to be stored during the subsequent storage process for the structured information extracted by the current extraction methods. For example, in practical applications, multiple structured storage systems are usually used to store structured information separately. However, the current extraction method directly uses neural networks and regular expressions to extract structured information from text. But there are usually a large number of data blocks in the text, and the structured information extracted from these data blocks may also need to be stored in different structured storage systems. Obviously, the current extraction method is difficult to achieve such classified storage.
[0021] In view of this, the embodiments of the present application provide a method and device for extracting structured information from text, which can be used to solve the problems in the prior art. For ease of understanding, an overall description of the embodiments of the present application can be provided here. The method can be applied to software products. For example, the method provided by the embodiments of the present application can be designed as an application program (i.e., software product), and then installed on an electronic device such as a user terminal or a server to apply the method provided by the embodiments of the present application.
[0022] For example, the software product can be installed on electronic devices at the user end, including mobile phones, tablet computers, computers and other electronic devices, and then the method provided in the embodiments of the present application can be executed; the software product can also be installed on electronic devices at the server end, including servers or server clusters, and then the method provided in the embodiments of the present application can be executed. In practical applications, the method can also be applied to hardware products. For example, dedicated hardware devices can be designed to implement the method provided in the embodiments of the present application.
[0023] Figure 1 A method for extracting structured information from text provided by an embodiment of the present invention. This method can be executed by an electronic device at the user end or by an electronic device at the server end. Here, taking the server (i.e., the electronic device at the server end) executing this method as an example, this method will be described. As Figure 1 shown, this method may include: Step S11: Determine the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted.
[0024] Among them, the text to be extracted can be any unstructured text (that is, the main data stored in this unstructured text is unstructured data). It is necessary to extract structured information from this unstructured text and store it in a structured storage system. Of course, the text to be extracted includes multiple data blocks. For example, the text to be extracted can be a contract text, and this contract text includes multiple chapters, and these different chapters can be used as different data blocks respectively; the text to be extracted can also be a news report, and different paragraphs of this news report can also be used as different data blocks; the text to be extracted can also be the text used to record business data during the daily business handling process of enterprise employees, and multiple data blocks can be divided according to the time when the business data is recorded. This attribute value degree measures the value of the text to be extracted from the level of the overall attributes of the text to be extracted.
[0025] It should be noted that in practical applications, in order to facilitate the preservation of the text, it is usually necessary to classify the text in combination with its content, so that the text has corresponding content classification labels. Among them, the content classification labels can include, for example, contract text, business data text, financial data text, annual, monthly and quarterly summary text, case data, meeting minutes, etc. Therefore, in this step S11, the content classification label of the text to be extracted can be obtained first, and then the attribute value degree of the text to be extracted can be determined. Specifically, the content classification label of the text to be extracted can be determined according to the text library stored in the text to be extracted or the file name, attribute value, etc. of the text to be extracted.
[0026] Of course, the specific method for determining the attribute value of the text to be extracted based on the content classification tags of the text to be extracted may include first obtaining the internal importance level of the text to be extracted and determining the importance score based on this internal importance level, then obtaining the network popularity of the content classification tag and determining the popularity score based on this network popularity, and then determining the attribute value of the text to be extracted based on this importance score and popularity score. Specifically, enterprises usually divide the importance levels of documents internally to facilitate the management of document access permissions (usually to prevent leakage). Therefore, the internal importance level of a document reflects the business value, relevance, or decision-making influence of the document within the enterprise, and further affects the degree of impact on the enterprise after its leakage.
[0027] For example, if financial data texts and contract texts are leaked, it will have an extremely serious impact on the operation of the enterprise. Therefore, their importance levels are usually extremely important; if business data texts and case data are leaked, it will also have a serious impact on the operation of the enterprise. Therefore, their importance levels are usually very important; the meeting minutes record important decisions and problem discussions within the enterprise. If leaked, it will have a certain impact on the operation of the enterprise, and its importance level is usually important.
[0028] Annual, monthly, and quarterly summary texts mainly summarize and analyze the work situation of the enterprise within a year, a month, or a quarter, including business results, work progress, existing problems, etc. Although it has certain value for the enterprise's own management and decision-making, the direct impact on the enterprise after leakage is relatively small, mainly possibly affecting the work rhythm and management efficiency within the enterprise. Therefore, the importance levels of annual, monthly, and quarterly summary texts are usually of general importance.
[0029] Of course, there are also some texts, including the enterprise's management system record texts, promotional brochures, etc. Since they usually need to be actively made public, for example, the enterprise actively makes public the promotional brochure for product promotion, their importance levels are usually relatively unimportant.
[0030] Specifically, for the texts with content classification tags such as the enterprise's contract texts, business data texts, financial data texts, annual, monthly, and quarterly summary texts, case data, meeting minutes, etc., their importance level divisions can be as shown in Table 1.
[0031] Table 1
[0032] In this application, the internal importance level of the text to be extracted can be obtained first. For example, the content classification label of the text to be extracted can be obtained, and then its internal importance level can be determined based on this content classification label (for example, the above Table 1 can be queried to determine its internal importance level). Then, the importance score can be determined based on this internal importance level. For example, in practical applications, different values can be assigned in advance for different internal importance levels, so that the corresponding importance score can be determined according to the internal importance level of the text to be extracted. Among them, since this importance score is determined based on the internal importance level of the text to be extracted, this importance score can be used to represent the importance of the text to be extracted within the enterprise.
[0033] For the specific method of obtaining the network popularity of the content classification label of the text to be extracted and determining the popularity score based on this network popularity, for example, the following method can be used to obtain the network popularity of the content classification label of the text to be extracted. For example, the network search volume, social media interaction volume, and news report mention volume of this content classification label can be obtained. Specifically, the network search volume, social media interaction volume, and news report mention volume of this content classification label can be obtained from major search engines, major social media interaction platforms, and major online news media platforms. Then, the following Formula 1 is used to calculate this network popularity: P 热度 =a×network search volume + b×social media interaction volume + c×news report mention volume Formula 1 In this Formula 1, P 热度 is the calculated network popularity; a, b, and c are preset weights respectively. Among them, the magnitudes of a, b, and c can usually be set according to actual needs. For example, if the enterprise business pays more attention to news reports, the value of c can be set relatively high. If the enterprise business pays more attention to social media interaction, the value of b can be set relatively high. Of course, if the enterprise's business is more related to the network search volume, the value of a can be set relatively high.
[0034] Specifically, for example, the content classification label of the text to be extracted is case data, that is, the case data of relevant cases related to this enterprise. For example, when this enterprise sends a lawsuit due to a contract dispute with other enterprises or individuals, the network search volume of relevant cases in major search engines such as Baidu and 360 Search (for example, Baidu Index can be used as this network search volume) can be obtained; the social media interaction volume of major social media platforms such as QQ, WeChat, DingTalk, and Weibo, and the news report mention volume of major online news media platforms such as Sina and Sohu. Then, substitute them into the above Formula 1 to calculate this network popularity.
[0035] After obtaining the network popularity, the popularity score can be further determined based on the network popularity. For example, the popularity score can be determined by Formula 2 as follows: H = (P 热度 - P min ) / (P max - P min ) Formula 2 In Formula 2, H is the calculated popularity score; P 热度 is the network popularity calculated by Formula 1 above; P max is the maximum value of the network popularity of this content classification label in each statistical period (such as one week as a statistical period) within the most recent year. For example, the network popularity of this content classification label can be counted weekly within the most recent year, and then the maximum value is selected as this P max ; P min is the minimum value of the network popularity of this content classification label in each statistical period within the most recent year. Similarly, the network popularity of this content classification label can be counted weekly within the most recent year, and then the minimum value is selected as this P min .
[0036] Therefore, after calculating the network popularity P 热度 of the content classification label of the text to be extracted through Formula 1 above, this P 热度 can be substituted into Formula 2 to further calculate the popularity score H of the content classification label of the text to be extracted. Obviously, this popularity score H can be used to represent the popularity of the content classification label of the text to be extracted in the online media.
[0037] After obtaining the importance score and the popularity score H through the above methods respectively, the attribute value of the text to be extracted can be further determined based on the importance score and the popularity score H. Among them, since the importance score is used to represent the importance of the text to be extracted within the enterprise, and the popularity score H is used to represent the popularity of the content classification label of the text to be extracted in the online media, the attribute value of the text to be extracted can measure the value of the text to be extracted from two perspectives of the importance within the enterprise and the popularity in the online media.
[0038] In practical applications, the importance score and the popularity score H can be substituted into Formula 3 as follows to calculate the attribute value of the text to be extracted: Formula 3 In Formula 3, P 属性 is the calculated attribute value; I is the importance score; H is the popularity score; both δ and γ are adjustment coefficients less than 1, and γ is greater than δ. Among them, the attribute value P calculated by Formula 3属性 From two perspectives, namely the importance within the enterprise and the popularity in online media, the value of the text to be extracted is measured. And since γ is greater than δ, the impact of the internal importance of the text to be extracted on P 属性 is greater. Among them, through Formula 3, the attribute value degree P on a 100-point scale can be calculated 属性 .
[0039] Step S12: For the target data block in the text to be extracted, determine the user behavior value degree of the target data block according to the user operation record of the target data block.
[0040] Among them, the target data block is specifically any data block in the text to be extracted. For example, in practical applications, structured information extraction can be performed on each data block in the text to be extracted respectively. At this time, each data block in the text to be extracted can be used as the target data block.
[0041] In practical applications, the user operation record may include the modification frequency of the target data block, the time interval between the current time and the most recent modification of the target data block, and the average viewing duration of the target data block. Among them, the modification frequency of the target data block refers to the frequency of modification of the target data block by the user. The specific modification methods may include adding content to the target data block, deleting content in the target data block, and modifying content in the target data block. Generally speaking, the higher the modification frequency, the higher the data activity of the target data block, and thus the higher its importance in the text to be extracted. In practical applications, the number of modifications of the target data block by the user can be recorded, and then the modification frequency of the target data block can be calculated by the ratio of the total duration after the generation of the target data block.
[0042] The time interval between the current time and the most recent modification of the target data block refers to the time interval between the current time and the most recent modification of the target data block. Specifically, the time of the most recent modification of the target data block can be determined, and then this time interval can be calculated. In practical applications, the larger this time interval is, the larger the time interval between the user's last modification of the target data block and the current time is, so the importance of the target data block in the text to be extracted is relatively lower. On the contrary, the smaller this time interval is, the higher the importance of the target data block in the text to be extracted is.
[0043] The average access duration of the target data block refers to the average value of the access duration of the target data block by the user each time. In practical applications, the access duration of the target data block by the user each time can be obtained by the duration of the user's stay after swiping to the target data block on the operation interface each time. Furthermore, the access durations of the target data block by the user each time are summed up, and the sum is divided by the number of accesses to obtain the average access duration of the target data block. Obviously, the greater the average access duration of the target data block, the higher the degree of attention of the user to the target data block.
[0044] Therefore, in this application, for the target data block in the text to be extracted, the specific method for determining the user behavior value degree of the target data block according to the user operation record of the target data block can be to determine the user behavior value degree of the target data block according to the modification frequency of the target data block, the time interval between the target data block and the most recent modification, and the average access duration of the target data block. Specifically, the data activity score can be determined according to the modification frequency of the target data block first, and the user attention score can be determined according to the average access duration of the target data block. The data freshness of the target data block is determined according to the time interval. Then, the data activity score and the user attention score are weighted and summed up, and the result of the weighted sum is multiplied by the data freshness to obtain the user behavior value degree. Among them, the weights of the data activity score and the user attention score can usually be values between 0 and 1, such as 0.3, 0.5, 0.75, etc.
[0045] Among them, the data freshness of the target data block can be calculated by Formula 4 shown as follows: Formula 4 In Formula 4, T is the data freshness of the target data block; t is the time interval; t 有效期 is the data validity period set for the target data block. In practical applications, it can usually be set as needed. For example, the data validity period t 有效期 of the target data block can be 1 year, half a year, etc. The data freshness T calculated by this formula 4 reflects the percentage of the remaining effective duration after the most recent modification of the target data block.
[0046] Of course, if the t is greater than t 有效期 , it means that the target data block has expired. At this time, the value of T can be 0. Therefore, in practical applications, it can be first determined whether the time interval is greater than the data validity period. If so, the value of the data freshness is 0. If not, the data freshness T can be calculated by this formula 4.
[0047] In addition, the method for determining the data activity score according to the modification frequency of the target data block may include obtaining the maximum value and the minimum value among the modification frequencies of each data block in the text to be extracted, and then calculating the data activity score by using the normalization calculation formula for the modification frequency of the target data block, the maximum value, and the minimum value. Among them, the normalization calculation formula may be, for example, F = 100×(f - f min ) / (f max - f min ). In this normalization calculation formula, F is the calculated data activity score; f is the modification frequency of the target data block; f min is the minimum value; f max is the maximum value.
[0048] The method for determining the user attention score according to the average viewing duration of the target data block may include obtaining the average viewing duration of the text to be extracted (i.e., the average value of the duration of each viewing of the text to be extracted by the user), and then determining the user attention score according to the average viewing duration of the text to be extracted and the average viewing duration of the target data block. Specifically, the user attention score can be determined by the formula . In this formula, D is the user attention score; min is the operator for taking the minimum value; d is the average viewing duration of the target data block; d total is the average viewing duration of the text to be extracted. Therefore, the meaning of this formula is that when the average viewing duration of the target data block is greater than the average viewing duration of the text to be extracted, the user attention score is determined to be 100 (with a full score of 100), and when the average viewing duration of the target data block is less than or equal to the average viewing duration of the text to be extracted, the value of D is .
[0049] In this way, the data activity score, the user attention score, and the data freshness can be calculated respectively through the above methods. Since all three are normalized to values on a 100-point scale, the data activity score and the user attention score can be further weighted and summed, and the result of the weighted sum is multiplied by the data freshness to obtain the user behavior value. Obviously, the user behavior value calculated by this method is different from the attribute value that measures the value of the text to be extracted from the overall attribute level of the text to be extracted. The user behavior value measures the value of the target data block from the perspective of various interaction behaviors of the user.
[0050] Step S13: Determine the comprehensive value score of the target data block according to the attribute value of the text to be extracted and the user behavior value of the target data block.
[0051] Among them, the attribute value degree obtained according to step S11 of the present application above is obviously a measure of the value of the text to be extracted from the perspective of the overall attributes of the text to be extracted, while the user behavior value degree obtained through step S12 above is a measure of the value of the target data block from the perspective of various interaction behaviors of the user. Therefore, in this step S13, the two can be further combined, that is, the comprehensive value score of the target data block can be determined according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block.
[0052] Specifically, usually, the attribute value degree and the user behavior value degree can be weighted and summed to calculate the comprehensive value score of the target data block. Among them, for the weights of the attribute value degree and the user behavior value, they can be determined according to actual needs, such as according to the business type of the enterprise itself and the content classification label of the text to be extracted. For example, if the enterprise pays more attention to sales-related businesses, the weight of the user behavior value degree can be increased. If the content classification label is contract text, financial data text, business data text, etc., the attribute value degree can be set relatively high. Of course, the two can also be set to the same weight.
[0053] Step S14: According to the comprehensive value score of the target data block, select a regular expression corresponding to the extraction granularity, and extract structured information from the target data block.
[0054] It should be emphasized that the present application sets regular expressions with multiple extraction granularities. Specifically, regular expressions with a coarse extraction granularity and regular expressions with a fine extraction granularity, as well as preset thresholds, can be preset. Among them, the regular expression with a coarse extraction granularity can be, for example, a look-ahead and look-behind assertion regular expression, which includes look-ahead and look-behind statements, and the keywords in the look-ahead and look-behind assertions can be one or more. This is the simplest regular expression.
[0055] The regular expression with a fine extraction granularity can be a multi-group chained regular expression or a nested regular expression. Among them, the multi-group chained regular expression refers to connecting multiple expressions through the use of a pipe symbol in the regular expression, and then taking the multiple expressions connected by the pipe symbol as a whole regular expression, that is, the multi-group chained regular expression. Therefore, on the one hand, the multi-group chained regular expression can achieve more fine-grained and more pattern-based matching of any one through each expression; the nested regular expression refers to nesting another regular expression in a regular expression, so as to also achieve a more complex and refined matching function.
[0056] Therefore, in this application, after obtaining the comprehensive value score of the target data block through the above step S13, for the specific implementation manner of step S14, it can first be determined whether the comprehensive value score of the target data block is greater than the preset threshold. At this time, when the comprehensive value score is greater than the preset threshold, it indicates that the target data block has a relatively high value for structured information extraction. Therefore, a regular expression with a fine extraction granularity can be selected to extract structured information from the target data block; or, when the comprehensive value score is less than or equal to the preset threshold, a regular expression with a coarse extraction granularity can be selected to extract structured information from the target data block.
[0057] The key point of step S14 in this application is to select a regular expression with a corresponding extraction granularity according to the comprehensive value score of the target data block, and then extract structured information from the target data block. However, for how to generate regular expressions and how to use regular expressions to extract structured information from data blocks, in practical applications, the methods in the prior art can be used to implement, and this is not limited here.
[0058] Step S15: Store the extracted structured information into the structured storage system corresponding to the comprehensive value score.
[0059] It should be further noted that this application includes multiple structured storage systems, and these different structured storage systems are respectively used to store structured information in different comprehensive value score intervals. For example, structured storage system 1 is used to store structured information in the comprehensive value score interval [P1, P2), structured storage system 2 is used to store structured information in the comprehensive value score interval [P2, P3), structured storage system 3 is used to store structured information in the comprehensive value score interval [P3, P4)...... In this way, after the structured information is extracted through the above step S14, in step S15, the structured storage system corresponding to the comprehensive value score can be determined according to the interval in which the obtained comprehensive value score falls, and then the extracted structured information can be stored in this structured storage system.
[0060] The method for extracting structured information from text provided by the embodiments of the present application includes: determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted; for the target data block in the text to be extracted, determining the user behavior value degree of the target data block according to the user operation record of the target data block; determining the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block; then, according to the comprehensive value score of the target data block, selecting a regular expression corresponding to the extraction granularity, extracting structured information from the target data block, and storing the extracted structured information in the structured storage system corresponding to the comprehensive value score. By calculating the comprehensive value score of the target data block in the text to be extracted, this method selects a regular expression corresponding to the extraction granularity, extracts structured information from the target data block, and stores the structured information in the structured storage system corresponding to the comprehensive value score. Therefore, even if there are multiple data blocks in the text to be extracted, this method can be used to extract structured information from each data block and store it in the corresponding structured storage system, thus solving the problems in the prior art.
[0061] Based on the same inventive concept as the method for extracting structured information from text provided by the embodiments of the present application, the embodiments of the present application can also provide a device for extracting structured information from text. For the content in the embodiments of this device, if there are unclear points, reference can be made to the relevant content in the above method embodiments. As Figure 2 shown in the specific structural schematic diagram of the device 20, the device 20 includes: a first determination unit 201, a second determination unit 202, a third determination unit 203, an extraction unit 204, and a storage unit 205, where: The first determination unit 201 is configured to determine the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted; The second determination unit 202 is configured to, for the target data block in the text to be extracted, determine the user behavior value degree of the target data block according to the user operation record of the target data block; where the target data block is specifically any one data block in the text to be extracted; The third determination unit 203 is configured to determine the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block; The extraction unit 204 is configured to select a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block, and extract structured information from the target data block; The storage unit 205 is configured to store the extracted structured information in the structured storage system corresponding to the comprehensive value score.
[0062] By using the device 20 provided in the embodiments of the present application, since the device 20 adopts the same inventive concept as the method provided in the embodiments of the present application, when the method can solve the problems in the prior art, the device 20 can also solve the problems in the prior art, and details are not described herein again.
[0063] The device 20 may further include a setting unit for presetting a regular expression for coarse extraction granularity, a regular expression for fine extraction granularity, and a preset threshold; and, According to the comprehensive value score of the target data block, select a regular expression with a corresponding extraction granularity, and extract structured information from the target data block, specifically including: Determine whether the comprehensive value score of the target data block is greater than the preset threshold; When the comprehensive value score is greater than the preset threshold, select a regular expression with a fine extraction granularity, and extract structured information from the target data block; or, When the comprehensive value score is less than or equal to the preset threshold, select a regular expression with a coarse extraction granularity, and extract structured information from the target data block.
[0064] Among them, according to the content classification label of the text to be extracted, determining the attribute value degree of the text to be extracted may specifically include: Obtain the internal importance level of the text to be extracted, and determine an importance score according to the internal importance level; wherein, the importance score is used to represent the importance of the text to be extracted within the enterprise; Obtain the network popularity of the content classification label, and determine a popularity score according to the network popularity; wherein, the popularity score is used to represent the popularity of the content classification label of the text to be extracted in network media; Determine the attribute value degree of the text to be extracted according to the importance score and the popularity score.
[0065] Among them, according to the importance score and the popularity score, determining the attribute value degree of the text to be extracted may specifically include calculating the attribute value degree through the following formula: ; Wherein: P_attribute is the attribute value degree; I is the importance score; H is the popularity score; both δ and γ are adjustment coefficients less than 1, and γ is greater than δ.
[0066] Among them, according to the user operation record of the target data block, determining the user behavior value degree of the target data block may specifically include: Determine the user behavior value of the target data block according to the modification frequency of the target data block, the time interval since the last modification of the target data block, and the average viewing duration of the target data block.
[0067] Among them, determining the user behavior value of the target data block according to the modification frequency of the target data block, the time interval since the last modification of the target data block, and the average viewing duration of the target data block may specifically include: Determine the data activity score according to the modification frequency, determine the user attention score according to the average viewing duration, and determine the data freshness of the target data block according to the time interval; Perform weighted summation on the data activity score and the user attention score, and multiply the result of the weighted summation by the data freshness to obtain the user behavior value.
[0068] Among them, determining the data activity score according to the modification frequency may specifically include: Obtain the maximum value and the minimum value among the modification frequencies of each data block in the text to be extracted; Calculate the data activity score by passing the modification frequency of the target data block, the maximum value, and the minimum value through a normalization calculation formula.
[0069] Among them, determining the data freshness of the target data block according to the time interval may specifically include: Judge whether the time interval is greater than the data validity period set for the target data block; If so, determine the data freshness as 0; or, If not, the data freshness can be calculated through the following formula; ; Among them, T is the data freshness; t is the time interval; t 有效期 is the data validity period set for the target data block.
[0070] Among them, determining the user attention score according to the average viewing duration may specifically include: Obtain the average viewing duration of the text to be extracted; Through the formula to determine the user attention score, where D is the user attention score; min is the operator for taking the minimum value; d is the average viewing duration of the target data block; d total is the average viewing duration of the text to be extracted.
[0071] Figure 3Illustrates a schematic diagram of the physical structure of an electronic device, as follows Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logical instructions in the memory 330 to execute the method provided in the embodiments of the present application for extracting structured information from text, including determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted, and for the target data block in the text to be extracted, determining the user behavior value degree of the target data block according to the user operation record of the target data block, determining the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block, and then selecting a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block, extracting structured information from the target data block, and then storing the extracted structured information in the structured storage system corresponding to the comprehensive value score. This method calculates the comprehensive value score of the target data block in the text to be extracted to select a regular expression corresponding to the extraction granularity, extracts structured information from the target data block, and stores the structured information in the structured storage system corresponding to the comprehensive value score. Therefore, even if there are multiple data blocks in the text to be extracted, the structured information in each data block can be extracted through this method and stored in the corresponding structured storage system, thus solving the problems in the prior art.
[0072] Among them, in practical applications, the electronic device may be an electronic device at the user end or an electronic device at the server end.
[0073] Obviously, since the processor 310 can call the logical instructions in the memory 330 to execute the method provided in the embodiments of the present application, it can also solve the problems in the prior art.
[0074] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0075] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for extracting structured information provided in the embodiments of the present application, including determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted, for the target data block in the text to be extracted, determining the user behavior value degree of the target data block according to the user operation record of the target data block, determining the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block, and then according to the comprehensive value score of the target data block, selecting a regular expression corresponding to the extraction granularity, extracting structured information from the target data block, and then storing the extracted structured information in the structured storage system corresponding to the comprehensive value score. This method calculates the comprehensive value score of the target data block in the text to be extracted to select a regular expression corresponding to the extraction granularity, extracts structured information from the target data block, and stores the structured information in the structured storage system corresponding to the comprehensive value score. Therefore, even if there are multiple data blocks in the text to be extracted, the structured information in each data block can be extracted by this method and stored in the corresponding structured storage system, thus solving the problems in the prior art.
[0076] Obviously, since the computer can execute the method provided in the embodiments of the present application when the computer program is executed by the processor, it can also solve the problems in the prior art.
[0077] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the embodiments of the present application is executed.
[0078] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for extracting structured information from text, characterized in that, The method includes: Determine the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted; For the target data block in the text to be extracted, determine the user behavior value degree of the target data block according to the user operation record of the target data block; wherein, the target data block is specifically any data block in the text to be extracted; Determine the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block; Select a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block, and extract structured information from the target data block; Store the extracted structured information in a structured storage system corresponding to the comprehensive value score.
2. The method according to claim 1, wherein The method further includes pre-setting a regular expression with a coarse extraction granularity, a regular expression with a fine extraction granularity, and a preset threshold; And, Selecting a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block and extracting structured information from the target data block specifically includes: Determine whether the comprehensive value score of the target data block is greater than the preset threshold; In the case where the comprehensive value score is greater than the preset threshold, select a regular expression with a fine extraction granularity and extract structured information from the target data block; or, In the case where the comprehensive value score is less than or equal to the preset threshold, select a regular expression with a coarse extraction granularity and extract structured information from the target data block.
3. The method according to claim 1, wherein Determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted specifically includes: Obtain the internal importance level of the text to be extracted, and determine an importance score according to the internal importance level; wherein, the importance score is used to represent the importance of the text to be extracted within the enterprise; Obtain the network popularity of the content classification label, and determine a popularity score according to the network popularity; wherein, the popularity score is used to represent the popularity of the content classification label of the text to be extracted in online media; Determine the attribute value degree of the text to be extracted according to the importance score and the popularity score.
4. The method according to claim 1, wherein Determining the user behavior value degree of the target data block according to the user operation record of the target data block specifically includes: Determine the user behavior value degree of the target data block according to the modification frequency of the target data block, the time interval between the target data block and the most recent modification, and the average viewing duration of the target data block.
5. The method according to claim 4, characterized in that, Determining the user behavior value degree of the target data block according to the modification frequency of the target data block, the time interval between the target data block and the most recent modification, and the average viewing duration of the target data block specifically includes: Determine a data activity score according to the modification frequency, determine a user attention score according to the average viewing duration, and determine the data freshness of the target data block according to the time interval; Perform a weighted sum of the data activity score and the user attention score, and multiply the result of the weighted sum by the data freshness to obtain the user behavior value degree.
6. The method according to claim 5, wherein Determine the data activity score according to the modification frequency, specifically including: Obtain the maximum and minimum values among the modification frequencies of each data block in the text to be extracted; Calculate the data activity score by using the modification frequency of the target data block, the maximum value, and the minimum value through a normalization calculation formula.
7. The method according to claim 5, characterized in that Determine the data freshness of the target data block according to the time interval, specifically including: Judge whether the time interval is greater than the data validity period set for the target data block; If so, determine the data freshness as 0; or, If not, calculate the data freshness through the following formula; ; Wherein, T is the data freshness; t is the time interval; t 有效期 is the data validity period set for the target data block.
8. The method according to claim 5, wherein Determine the user attention score according to the average viewing duration, specifically including: Obtain the average viewing duration of the text to be extracted; Determine the user importance score through the formula where D is the user importance score; min is the operator for taking the minimum value; d is the average access duration of the target data block; and d total is the average access duration of the text to be extracted.
9. An apparatus for extracting structured information from text, characterized in that, including: A first determination unit for determining the attribute value degree of the text to be extracted according to the content classification label of the text to be extracted; A second determination unit for determining the user behavior value degree of the target data block in the text to be extracted according to the user operation record of the target data block; wherein, the target data block is specifically any data block in the text to be extracted; A third determination unit for determining the comprehensive value score of the target data block according to the attribute value degree of the text to be extracted and the user behavior value degree of the target data block; An extraction unit for selecting a regular expression corresponding to the extraction granularity according to the comprehensive value score of the target data block and extracting structured information from the target data block; A storage unit for storing the extracted structured information into a structured storage system corresponding to the comprehensive value score.
Citation Information
Patent Citations
A method, device, and storage medium for accurately extracting structured information from complex web pages.
CN113254751B
Text structured information extraction method based on neural network and related equipment thereof
CN113987125A
Scientific and technical journal article word document structural processing method and device
CN108153717A
Webpage data structured extraction method
CN111523303A
Subject parameter information extraction method and device, storage medium and electronic equipment
CN113590655A