Data recommendation method based on large model

By determining the search status based on the reference quality index and text correlation coefficient of the search text set, and extracting the recommended text content using appropriate analysis methods, the problem of semantic breaks in the recommended text in the prior art is solved, and the effectiveness and coherence of the recommended text content is improved.

CN119719353BActive Publication Date: 2025-05-23ZHONGHE YUNKE INFORMATION TECH GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510205670.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-23
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art has failed to determine a targeted extraction method based on the correlation of recommended segments in the search text, resulting in semantic breaks easily in the recommended text and low effectiveness.

Method used

By obtaining the search keywords of the target problem text, obtaining the search text collection, and determining the search status based on the reference quality index and text correlation coefficient of the search text collection, the recommended text content is extracted using paragraph analysis or combination analysis to ensure that the extraction method meets the correlation of the recommended segments.

Benefits of technology

It improves the effectiveness and coherence of the recommended text content, avoids semantic breakage, and enhances the information relevance of the recommended text content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719353B_ABST
    Figure CN119719353B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of text data analysis, and in particular to a data recommendation method based on a large model, comprising obtaining retrieval keywords of a target question text, obtaining a retrieval text set of the target question text based on the retrieval keywords and related keywords; performing retrieval association evaluation on the retrieval text set to determine the retrieval status of the retrieval text set, and determining a text processing strategy according to the retrieval status; when a paragraph analysis method is adopted, determining a paragraph extraction method corresponding to each text paragraph according to a reference extraction coefficient of each text paragraph; when a combined analysis method is adopted, obtaining key analysis texts and key keywords, and determining matching text paragraphs according to a matching keyword proportion and matching effectiveness; the present invention improves the effectiveness of the obtained recommended text content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text data analysis, and in particular to a data recommendation method based on a large model. Background Art

[0002] Currently, the quality of text generation results obtained by large-scale language models depends on the quality of text retrieval and the quality of the extracted recommended text content. In the text generation process, it is necessary to extract recommended segments for the retrieved text. During the extraction process, the correlation between the recommended segments in the text affects the effectiveness and readability of the obtained recommended text content. Therefore, how to avoid the low effectiveness of the recommended text content caused by excessive segmentation in the process of extracting recommended segments is a problem that needs to be urgently solved by technical personnel in this field.

[0003] Chinese patent publication number CN117195890A discloses a text recommendation method based on machine learning, which belongs to the field of semantic extraction technology. In this invention, each keyword in the user information string is arranged and combined to obtain different information sequences. This invention measures the fit score of an information sequence according to the number of keywords contained in each information sequence and the weight of the keywords contained. In the invention, the semantic features of each information sequence are identified by a machine learning model to achieve further semantic extraction of each information sequence, calculate the matching degree of the semantic features of each information sequence with the text to be recommended, and then comprehensively calculate the fit score of the information sequence and the user information string, calculate the recommendation score of the text to be recommended, and recommend all relevant texts. However, the above scheme has the following problems: it fails to determine the targeted extraction method of the text to be recommended according to the association of each text to be recommended in the long text to which it belongs, resulting in the easy existence of semantic breaks in the content of the recommended text obtained, which in turn leads to the low effectiveness of the recommended text content. Summary of the invention

[0004] To this end, the present invention provides a data recommendation method based on a large model to overcome the problem in the prior art that a targeted recommended paragraph extraction method is unable to be determined according to the association of each recommended paragraph in the search text to which it belongs, resulting in semantic breaks in the obtained recommended text, thereby leading to low effectiveness of the recommended text content.

[0005] To achieve the above object, the present invention provides a data recommendation method based on a large model, comprising:

[0006] Obtaining search keywords of the target question text, obtaining a search text set of the target question text based on the search keywords and related keywords, and performing format conversion on the target search text contained therein;

[0007] Conduct retrieval association evaluation on the retrieval text set to determine the retrieval status of the retrieval text set, and determine the text processing strategy based on the retrieval status. The text processing strategy is to extract the recommended text content by using a paragraph analysis method or a combination analysis method;

[0008] When the paragraph analysis method is adopted, the paragraph extraction method corresponding to each text paragraph is determined according to the reference extraction coefficient of each text paragraph. The paragraph extraction method is to perform segmentation compensation for the text paragraph according to the paragraph association relationship, or to determine the recommended text paragraph of the text paragraph according to the paragraph extraction coefficient;

[0009] When using the combined analysis method, the key analysis text and key keywords of the search text set are obtained, and the matching text segments are determined based on the matching keyword proportion and matching effectiveness;

[0010] Send the recommended text content to the user.

[0011] Furthermore, under the high-quality search condition, the search status of the search text set is determined according to the reference quality index and the text correlation coefficient.

[0012] If the reference quality index of the search text set is greater than the preset reference quality index, it is determined that the search text set is in a first preset search state;

[0013] If the reference quality index of the search text set is less than or equal to the preset reference quality index and the text correlation coefficient is greater than the preset text correlation coefficient, it is determined that the search text set is in the second preset search state;

[0014] The high-quality search condition is that the proportion of high-quality texts in the search text set is greater than a preset proportion of high-quality texts.

[0015] Furthermore, the process of performing search relevance evaluation on the search text set includes:

[0016] Determine the search quality index of each target search text in the search text set according to the keyword coverage and the proportion of related texts, and record the target search text with a search quality index greater than the preset search quality index as a high-quality search text;

[0017] The text correlation coefficient of the retrieval text set is determined according to the effective keywords of each target retrieval text, and the average value of the retrieval quality index of each target retrieval text is recorded as the reference quality index of the retrieval text set.

[0018] Further, determining a text processing strategy according to a retrieval status of the retrieval text set;

[0019] If the search text set is in the first preset search state, extracting recommended text content from each target search text included in the search text set by using a paragraph analysis method;

[0020] If the search text set is in the second preset search state, a combined analysis method is used to extract recommended text content from each target search text included in the search text set.

[0021] Furthermore, when extracting the recommended text content from a target search text by using a paragraph analysis method, the paragraph extraction coefficient of each target text paragraph in the target search text is determined according to the reference association frequency of each effective keyword and the search level difference coefficient;

[0022] The paragraph extraction coefficient is positively correlated with the reference association frequency of the effective keywords in the target text paragraph, and the paragraph extraction coefficient is negatively correlated with the retrieval level difference coefficient of the effective keywords in the target text paragraph.

[0023] Further, determining the paragraph extraction method corresponding to each text paragraph according to the reference extraction coefficient of each text paragraph;

[0024] If the reference extraction coefficient of a text paragraph is greater than the preset reference extraction coefficient, segmentation compensation is performed on the text paragraph according to the paragraph association relationship;

[0025] If the reference extraction coefficient of a text paragraph is less than or equal to the preset reference extraction coefficient, the recommended text of the text paragraph is determined according to the paragraph extraction coefficient.

[0026] Furthermore, when segmenting and compensating a text paragraph according to the paragraph association relationship, the content association coefficient of each text paragraph to be analyzed is determined according to the keyword overlap degree and the keyword matching degree, and the retained paragraph of each text paragraph is determined according to the content association coefficient and the text connection coefficient, and the retained paragraph and the recommended text paragraph are recorded as the recommended text of the text paragraph;

[0027] The recommended text segment is a target text segment whose segment extraction coefficient is greater than a preset segment extraction coefficient, and the to-be-analyzed text segment is a target text segment whose segment extraction coefficient is less than or equal to the preset segment extraction coefficient.

[0028] Furthermore, when extracting recommended text content from a search text set by using a combined analysis method, the target search text with the largest search quality index in the search text set is recorded as the key analysis text, and the key keywords are determined according to the application reference values ​​and reference association frequencies of the effective keywords contained in the key analysis text;

[0029] The target search text containing the key keywords is recorded as the relevant analysis text.

[0030] Furthermore, matching paragraph analysis is performed on each text paragraph in the key analysis text. For a single text paragraph, the matching paragraph analysis process includes:

[0031] Determine the recommended text paragraphs of the text paragraph and the relevant text paragraphs of each related analysis text according to the key keywords;

[0032] The matching text segment of the text paragraph is determined according to the matching richness and matching effectiveness of each related text segment.

[0033] Furthermore, if the matching text segments are determined for each text paragraph in the key analysis text, the matching segment combination is determined according to the segment overlap coefficient, and the matching key keywords of each matching segment combination are determined according to the distribution reference value and the search relevance;

[0034] The segment overlap coefficients between the matching text segments in any matching segment combination are all greater than the preset segment overlap coefficient.

[0035] Compared with the prior art, the beneficial effect of the present invention lies in that the technical solution of the present invention determines the retrieval status of the retrieval text set according to the reference quality index and the text correlation coefficient, and determines the text processing strategy according to the retrieval status of the retrieval text set, so that the method adopted for extracting the recommended text content for the target retrieval text is more in line with the association of the recommended paragraphs in the target retrieval text, avoiding excessive segmentation of the recommended paragraphs during the extraction process, and the present invention improves the effectiveness of the obtained recommended text content.

[0036] Furthermore, the present invention determines the retrieval status of the retrieval text set based on the reference quality index and the text correlation coefficient, which is used to characterize the distribution of recommended segments in the retrieval text set in each target retrieval text and the degree of correlation between the target retrieval texts. The retrieval status of the retrieval text set is determined in this way, making the determination results of subsequent text processing strategies more accurate and improving the effectiveness of extracting recommended segments.

[0037] Furthermore, in the present invention, for the retrieval text set in the first preset retrieval state, it is determined whether to perform segmentation compensation on the text paragraph when extracting the target text segment according to the reference extraction coefficient of each text paragraph. If the reference extraction coefficient of a text paragraph is greater than the preset reference extraction coefficient, it indicates that the recommended text segments account for a large proportion in the text paragraph, and there is a large correlation between the recommended text segments. It is necessary to retain some segments of the text paragraph during the extraction process to avoid semantic breaks in the recommended text content, so as to ensure the coherence and effectiveness of the obtained recommended text content.

[0038] Furthermore, in the present invention, for a retrieval text set in a second preset retrieval state, the key analysis text and key keywords of the retrieval text set are determined according to the retrieval quality index, and the relevant analysis texts in the retrieval text set are analyzed based on the key analysis text and its key keywords to obtain matching text segments, and the matching text segments are combined with the recommended text segments in the key analysis text, thereby improving the relevance and effectiveness of the information in the obtained recommended text content. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A schematic diagram of a data recommendation method based on a large model of the present invention;

[0040] Figure 2 A flowchart of the present invention for determining the retrieval status of a retrieval text set according to a reference quality index and a text correlation coefficient;

[0041] Figure 3 A flowchart of the present invention for determining a text processing strategy according to a search state of a search text set;

[0042] Figure 4 The present invention is a flow chart of a method for determining a paragraph extraction method according to a reference extraction coefficient of each text paragraph. DETAILED DESCRIPTION

[0043] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0045] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0046] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0047] See also Figures 1 to 4 As shown, the present invention provides a data recommendation method based on a large model, comprising:

[0048] Obtaining search keywords of the target question text, obtaining a search text set of the target question text based on the search keywords and related keywords, and performing format conversion on the target search text contained therein;

[0049] Conduct retrieval association evaluation on the retrieval text set to determine the retrieval status of the retrieval text set, and determine the text processing strategy based on the retrieval status. The text processing strategy is to extract the recommended text content by using a paragraph analysis method or a combination analysis method;

[0050] When the paragraph analysis method is adopted, the paragraph extraction method corresponding to each text paragraph is determined according to the reference extraction coefficient of each text paragraph. The paragraph extraction method is to perform segmentation compensation for the text paragraph according to the paragraph association relationship, or to determine the recommended text paragraph of the text paragraph according to the paragraph extraction coefficient;

[0051] When using the combined analysis method, the key analysis text and key keywords of the search text set are obtained, and the matching text segments are determined based on the matching keyword proportion and matching effectiveness;

[0052] Send the recommended text content to the user.

[0053] Among them, the present invention is provided with a keyword library, and the keywords contained therein are divided and the hierarchical coefficients are set accordingly, the keywords existing in the target question text are obtained, recorded as search keywords, and the keywords in the keyword library whose hierarchical difference value with the search keywords is less than the preset hierarchical difference value are recorded as related keywords, and the hierarchical difference value is the absolute value of the difference between the hierarchical coefficients of two valid keywords. The value of the preset hierarchical difference value can be determined by the user according to the actual working scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirement for the validity of the recommended text content, the smaller the value of the preset hierarchical difference value. A method for determining the preset hierarchical difference value is provided, and the average value of the hierarchical difference values ​​of each associated combination in the text recommendation record that meets the user's requirement for the validity of the recommended text content is recorded as the preset hierarchical difference value;

[0054] Based on the search keywords and related keywords, different long texts obtained are recorded as target search texts, and each target search text is divided into a search text set, wherein the target search texts obtained have different formats, and the format types of the target search texts include but are not limited to: pdf, txt and html, and the target search texts of different formats are converted into markdown format, and each target text segment in the target search text after the format conversion corresponds to various levels of titles, and the set of target text segments with the same corresponding various levels of titles is recorded as a text paragraph, and any target text segment is a complete sentence in the target search text;

[0055] How to divide and set the hierarchical coefficient for each keyword in the keyword library is easy for technical personnel in this field to understand. The specific execution method is not limited here. An example of dividing and setting the hierarchical coefficient for each keyword in the keyword library is provided: if the keywords in the keyword library are electrical appliances, smart phones, computer equipment, household appliances, refrigerators, folding screen mobile phones, curved screen mobile phones, household appliance brand A, mobile phone brand B, color, and storage capacity, the hierarchical coefficient of "electrical appliances" is set to 1, and "smart phones", "computer equipment" and "household appliances" are recorded as subcategories of "electrical appliances", and their hierarchical coefficients are set to 2, "refrigerator" is recorded as a subcategory of "household appliances", and its hierarchical coefficient is set to 3, "household appliance brand A" is recorded as a subcategory of "refrigerator", and its hierarchical coefficient is set to 4, "mobile phone brand B" is recorded as a subcategory of "smart phones", and its hierarchical coefficient is set to 3, "folding screen mobile phones", "curved screen mobile phones", "color" and "storage capacity" are recorded as subcategories of "mobile phone brand B", and their hierarchical coefficients are set to 4.

[0056] The present invention uses several text recommendation records, and any one of the text recommendation records records the proportion of high-quality texts, reference quality index, text correlation coefficient, retrieval quality index, hierarchical difference value, reference extraction coefficient, paragraph extraction coefficient, content correlation coefficient, matching coefficient, paragraph overlap coefficient and key analysis coefficient in the process of extracting recommended text content for a target problem text at least once, and each text recommendation record corresponds to a qualified mark, which records whether the user's requirements for the effectiveness of the recommended text content meet the user's needs.

[0057] Specifically, under high-quality search conditions, the search status of the search text set is determined based on the reference quality index and the text relevance coefficient.

[0058] If the reference quality index of the search text set is greater than the preset reference quality index, it is determined that the search text set is in a first preset search state;

[0059] If the reference quality index of the search text set is less than or equal to the preset reference quality index and the text correlation coefficient is greater than the preset text correlation coefficient, it is determined that the search text set is in the second preset search state;

[0060] The high-quality search condition is that the proportion of high-quality texts in the search text set is greater than a preset proportion of high-quality texts.

[0061] Among them, the high-quality text ratio = the number of high-quality search texts in the search text set / the number of target search texts in the search text set. The value of the preset high-quality text ratio can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirements for the effectiveness of the recommended text content, the larger the value of the preset high-quality text ratio. A method for determining the value of the preset high-quality text ratio is provided. The average value of the high-quality text ratio of each search text set that performs recommended text content extraction in the text recommendation record that meets the user's requirements for the effectiveness of the recommended text content is recorded as the preset high-quality text ratio. A value of the preset high-quality text ratio is provided, and the value of the preset high-quality text ratio is 0.6;

[0062] If the reference quality index of the search text set is less than or equal to the preset reference quality index and the text correlation coefficient is less than or equal to the preset text correlation coefficient, the search status of the search text set is not determined, and only the set of target text segments with valid keywords is recorded as the recommended text content of the search text set;

[0063] The values ​​of the preset reference quality index and the preset text correlation coefficient can be determined by the user according to the actual work scenario. For example, the user can set them according to the text recommendation record, and a method for determining the value of the preset reference quality index is provided, in which the text recommendation record that performs segmentation compensation in the process of extracting the recommended paragraph is recorded as a first-class reference record, and the minimum value of the reference quality index in the first-class reference record that meets the user's requirements for the effectiveness of the recommended text content is recorded as the preset reference quality index. A method for determining the value of the preset text correlation coefficient is provided, in which the text recommendation record that extracts the recommended text content using a combined analysis method is recorded as a second-class reference record, and the minimum value of the text correlation coefficient in the second-class reference record that meets the user's requirements for the effectiveness of the recommended text content is recorded as the preset text correlation coefficient.

[0064] Specifically, the process of conducting retrieval relevance evaluation on a retrieval text set includes:

[0065] Determine the search quality index of each target search text in the search text set according to the keyword coverage and the proportion of related texts, and record the target search text with a search quality index greater than the preset search quality index as a high-quality search text;

[0066] The text correlation coefficient of the retrieval text set is determined according to the effective keywords of each target retrieval text, and the average value of the retrieval quality index of each target retrieval text is recorded as the reference quality index of the retrieval text set.

[0067] Among them, for a single target search text, the search quality index is the sum of the products of the keyword coverage and the relevant text ratio and their corresponding impact coefficients. The search keywords and relevant keywords contained in the target search text are recorded as the effective keywords of the target search text. Keyword coverage = the number of effective keywords contained in the target search text / the sum of the number of search keywords and relevant keywords in the target question text. Relevant text ratio = the sum of the number of words in each relevant paragraph in the target search text / the total number of words contained in the target search text. The relevant paragraph is the target text paragraph where the effective keywords of the target search text exist. Provide a keyword coverage and the impact coefficient corresponding to the relevant text ratio. The impact coefficient corresponding to the keyword coverage is 0.5, and the impact coefficient corresponding to the relevant text ratio is 0.5.

[0068] The value of the preset retrieval quality index can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirement for the effectiveness of the recommended text content, the larger the value of the preset retrieval quality index. A method for setting the value of the preset retrieval quality index is provided, and the minimum value of the retrieval quality index of high-quality retrieval texts in the text recommendation record that meets the user's requirement for the effectiveness of the recommended text content is recorded as the preset retrieval quality index; the text correlation coefficient is the number of associated combinations existing in the retrieval text set. If the hierarchical difference value between two valid keywords respectively existing in two target retrieval texts is less than the preset hierarchical difference value, the above two valid keywords are recorded as an associated combination.

[0069] Specifically, a text processing strategy is determined according to the retrieval status of the retrieval text collection;

[0070] If the search text set is in the first preset search state, extracting recommended text content from each target search text included in the search text set by using a paragraph analysis method;

[0071] If the search text set is in the second preset search state, a combined analysis method is used to extract recommended text content from each target search text included in the search text set.

[0072] Among them, if the search text set is in the first preset search state, the set of recommended texts of each text paragraph is recorded as the recommended text content of the search text set; if the search text set is in the second preset search state, the set of recommended text segments of each text paragraph in the key analysis text and their matching segment combinations are recorded as the recommended text content of the search text set.

[0073] Specifically, when extracting the recommended text content from a target search text using a paragraph analysis method, the paragraph extraction coefficient of each target text paragraph in the target search text is determined according to the reference association frequency of each effective keyword and the search level difference coefficient;

[0074] The paragraph extraction coefficient is positively correlated with the reference association frequency of the effective keywords in the target text paragraph, and the paragraph extraction coefficient is negatively correlated with the retrieval level difference coefficient of the effective keywords in the target text paragraph.

[0075] Among them, for a single target text segment, the segment extraction coefficient is the sum of the association parameters of each valid keyword contained in the target text segment, the association parameter of any valid keyword = ln (reference association frequency / retrieval level difference coefficient), the reference association frequency = the number of key recommended texts containing the valid keyword / the number of key recommended texts, the key recommended text is the recommended text content containing the retrieval keyword of the target question text currently being recommended text content acquired in the text recommendation record, if the valid keyword is a retrieval keyword, the reference association frequency of the valid keyword is recorded as 1, and the retrieval level difference coefficient is the minimum value of the level difference value between the valid keyword and each retrieval keyword.

[0076] Specifically, the paragraph extraction method corresponding to each text paragraph is determined according to the reference extraction coefficient of each text paragraph;

[0077] If the reference extraction coefficient of a text paragraph is greater than the preset reference extraction coefficient, segmentation compensation is performed on the text paragraph according to the paragraph association relationship;

[0078] If the reference extraction coefficient of a text paragraph is less than or equal to the preset reference extraction coefficient, the recommended text of the text paragraph is determined according to the paragraph extraction coefficient.

[0079] Among them, for a single text paragraph, the reference extraction coefficient is the average value of the segment extraction coefficients of the target text segments contained in the text paragraph. The value of the preset reference extraction coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirements for the validity of the recommended text content, the smaller the value of the preset reference extraction coefficient. A method for setting the value of the preset reference extraction coefficient is provided, and the minimum value of the reference extraction coefficient of each text paragraph for segmentation compensation in a class of reference records that meet the user's requirements for the validity of the recommended text content is recorded as the preset reference extraction coefficient;

[0080] If the reference extraction coefficient of a text paragraph is less than or equal to the preset reference extraction coefficient, the target text paragraph whose paragraph extraction coefficient is greater than the preset paragraph extraction coefficient is recorded as the recommended text paragraph of the text paragraph, and the set of recommended text paragraphs is recorded as the recommended text of the text paragraph. The value of the preset paragraph extraction coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirement for the validity of the recommended text content, the larger the value of the preset paragraph extraction coefficient. A method for determining the value of the preset paragraph extraction coefficient is provided, and the minimum value of the paragraph extraction coefficient of the recommended text paragraphs in a class of reference records that meet the user's requirement for the validity of the recommended text content is recorded as the preset paragraph extraction coefficient.

[0081] Specifically, when segmenting and compensating a text paragraph according to the paragraph association relationship, the content association coefficient of each text paragraph to be analyzed is determined according to the keyword overlap degree and the keyword matching degree, and the retained paragraph of each text paragraph is determined according to the content association coefficient and the text connection coefficient, and the retained paragraph and the recommended text paragraph are recorded as the recommended text of the text paragraph;

[0082] The recommended text segment is a target text segment whose segment extraction coefficient is greater than a preset segment extraction coefficient, and the to-be-analyzed text segment is a target text segment whose segment extraction coefficient is less than or equal to the preset segment extraction coefficient.

[0083] Among them, for a single text segment to be analyzed, the content relevance coefficient is the sum of the products of the keyword overlap and the keyword match and their corresponding evaluation influence coefficients, keyword overlap = the number of overlapping keywords in the text to be analyzed / the sum of the number of valid keywords contained in each recommended text segment, overlapping keywords are valid keywords contained in both the text to be analyzed and the recommended text segment, and the keyword match is the average value of the retrieval level difference coefficients of each overlapping keyword in the text to be analyzed. The user can set the values ​​of the evaluation influence coefficients corresponding to the keyword overlap and the keyword match according to the internship work scenario, and provide a value of the evaluation influence coefficient corresponding to the keyword overlap and the keyword match. The value of the evaluation influence coefficient corresponding to the keyword overlap is 0.7, and the value of the evaluation influence coefficient corresponding to the keyword match is 0.3. The text connection coefficient is the number of recommended text segments that have a logical relationship with the text segment to be analyzed. The logical relationship between the target text segments in the present invention includes but is not limited to: cause and effect, comparison and contrast, and parallel relationship;

[0084] If the content association coefficient of a text segment to be analyzed is greater than the preset content association coefficient or the text connection coefficient is greater than the preset text connection coefficient, the text segment to be analyzed is recorded as a reserved segment. The values ​​of the preset content association coefficient and the preset text connection coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirement for the effectiveness of the recommended text content, the larger the value of the preset content association coefficient and the larger the value of the preset text connection coefficient. A method for determining the value of the preset content association coefficient is provided, and the average value of the content association coefficients of each reserved text in a class of reference records that meet the user's requirements for the effectiveness of the recommended text content is recorded as the preset content association coefficient. A method for determining the value of the preset text connection coefficient is provided, and the average value of the text connection coefficients of each reserved text in a class of reference records that meet the user's requirements for the effectiveness of the recommended text content is recorded as the preset text connection coefficient.

[0085] Specifically, when extracting recommended text content from a search text set using a combined analysis method, the target search text with the largest search quality index in the search text set is recorded as the key analysis text, and the key keywords are determined based on the application reference values ​​and reference association frequencies of the effective keywords contained in the key analysis text;

[0086] The target search text containing the key keywords is recorded as the relevant analysis text.

[0087] Among them, the key analysis coefficient of each effective keyword is determined according to the application reference value and the reference association frequency, and the effective keyword in the key analysis text whose key analysis coefficient is greater than the preset key analysis coefficient is recorded as the key keyword. For a single effective keyword, the key analysis coefficient = ln (application reference value × reference association frequency), and the application reference value is the number of target text segments in the key analysis text that contain the effective keyword. The value of the preset key analysis coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirement for the effectiveness of the recommended text content, the larger the value of the preset key analysis coefficient. A method for determining the value of the preset key analysis coefficient is provided, and the minimum value of the key analysis coefficient of the key keyword in the second category of reference records that meets the user's requirements for the recommended text content is recorded as the preset key analysis coefficient.

[0088] Specifically, matching paragraph analysis is performed on each text paragraph in the key analysis text. For a single text paragraph, the matching paragraph analysis process includes:

[0089] Determine the recommended text paragraphs of the text paragraph and the relevant text paragraphs of each related analysis text according to the key keywords;

[0090] The matching text segment of the text paragraph is determined according to the matching richness and matching effectiveness of each related text segment.

[0091] Among them, for a single text paragraph, the target text segment containing the key keyword in the target search text other than the key analysis text in the search text set is recorded as the relevant text segment of the text paragraph, and the target text segment with the key keyword in the text paragraph is recorded as the recommended text segment of the text paragraph. For a single relevant text segment, the matching coefficient of the relevant text segment is determined according to the matching richness and matching effectiveness. The matching coefficient is the product of the matching richness and the matching effectiveness. The matching richness is the number of different key keywords contained in the relevant text segment. The matching effectiveness is the average value of the retrieval level difference coefficients of the key keywords contained in the relevant text segment.

[0092] The relevant text segments whose matching coefficients are greater than the preset matching coefficients are recorded as matching text segments. The value of the preset matching coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirements for the effectiveness of the recommended text content, the larger the value of the preset matching coefficient. A method for determining the value of the preset matching coefficient is provided, and the minimum value of the matching coefficients of each matching text segment in the two types of reference records that meet the user's requirements for the recommended text content is recorded as the preset matching coefficient.

[0093] Specifically, if the matching text segments of each text paragraph in the key analysis text are determined, the matching segment combination is determined according to the segment overlap coefficient, and the matching key keywords of each matching segment combination are determined according to the distribution reference value and the search relevance;

[0094] The segment overlap coefficients between the matching text segments in any matching segment combination are all greater than the preset segment overlap coefficient.

[0095] Wherein, for any matching paragraph combination, the paragraph overlap coefficient is the number of overlapping phrases in the matching paragraph combination. For a single key keyword, if the number of matching text paragraphs of the key keyword in the matching paragraph combination is greater than the preset number of paragraphs, the key keyword is recorded as the overlapping phrase of the matching paragraph combination. The value of the preset number of paragraphs can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirements for the effectiveness of the recommended text content, the larger the value of the preset number of paragraphs. A value of the preset number of paragraphs is provided, and the value of the preset number of paragraphs is 50% of the number of matching text paragraphs contained in the matching paragraph combination.

[0096] The value of the preset segment overlap coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to the text recommendation record. The higher the user's requirements for the effectiveness of the recommended text content, the larger the value of the preset segment overlap coefficient. A method for determining the value of the preset segment overlap coefficient is provided. The minimum value of the segment overlap coefficient between the matching text segments in each matching segment combination in the two types of reference records that meet the user's requirements for the recommended text content is recorded as the preset segment overlap coefficient.

[0097] For any matching paragraph combination, the distribution reference value and search relevance of the key keywords it contains are tested to determine the matching priority coefficient of each key keyword. For a single key keyword, the matching priority coefficient = ln (distribution reference value × search relevance), the distribution reference value is the number of matching text paragraphs of the key keyword in the matching paragraph combination, and the search relevance is the average of the hierarchical difference values ​​between the key keyword and each search keyword. The key keyword with the largest matching priority coefficient is selected and recorded as the matching key keyword of the matching paragraph combination.

[0098] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data recommendation method based on a large model, characterized in that: include: Obtaining search keywords of the target question text, obtaining a search text set of the target question text based on the search keywords and related keywords, and performing format conversion on the target search text contained in the search text set; Conduct retrieval association evaluation on the retrieval text set to determine the retrieval status of the retrieval text set, and determine the text processing strategy based on the retrieval status. The text processing strategy is to extract the recommended text content by using a paragraph analysis method or a combination analysis method; When the paragraph analysis method is adopted, the paragraph extraction method corresponding to each text paragraph is determined according to the reference extraction coefficient of each text paragraph. The paragraph extraction method is to perform segmentation compensation for the text paragraph according to the paragraph association relationship, or to determine the recommended text paragraph of the text paragraph according to the paragraph extraction coefficient; When using the combined analysis method, the key analysis text and key keywords of the search text set are obtained, and the matching text segments are determined based on the matching keyword proportion and matching effectiveness; Send the recommended text content to the user; Under high-quality search conditions, the search status of the search text set is determined based on the reference quality index and text relevance coefficient. If the reference quality index of the search text set is greater than the preset reference quality index, it is determined that the search text set is in a first preset search state; If the reference quality index of the search text set is less than or equal to the preset reference quality index and the text correlation coefficient is greater than the preset text correlation coefficient, it is determined that the search text set is in the second preset search state; The high-quality search condition is that the proportion of high-quality texts in the search text set is greater than a preset proportion of high-quality texts; The process of evaluating the retrieval relevance of a retrieval text set includes: Determine the search quality index of each target search text in the search text set according to the keyword coverage and the proportion of related texts, and record the target search text with a search quality index greater than the preset search quality index as a high-quality search text; The text correlation coefficient of the retrieval text set is determined according to the effective keywords of each target retrieval text, and the average value of the retrieval quality index of each target retrieval text is recorded as the reference quality index of the retrieval text set.

2. The data recommendation method based on a large model according to claim 1, characterized in that: Determining a text processing strategy based on a retrieval status of a retrieval text collection; If the search text set is in the first preset search state, extracting recommended text content from each target search text included in the search text set by using a paragraph analysis method; If the search text set is in the second preset search state, a combined analysis method is used to extract recommended text content from each target search text included in the search text set.

3. The data recommendation method based on a large model according to claim 2, characterized in that: When extracting the recommended text content from a target search text by using the paragraph analysis method, the paragraph extraction coefficient of each target text paragraph in the target search text is determined according to the reference association frequency of each effective keyword and the search level difference coefficient; The paragraph extraction coefficient is positively correlated with the reference association frequency of the effective keywords in the target text paragraph, and the paragraph extraction coefficient is negatively correlated with the retrieval level difference coefficient of the effective keywords in the target text paragraph.

4. The data recommendation method based on a large model according to claim 3, characterized in that: Determine the paragraph extraction method corresponding to each text paragraph according to the reference extraction coefficient of each text paragraph; If the reference extraction coefficient of a text paragraph is greater than the preset reference extraction coefficient, segmentation compensation is performed on the text paragraph according to the paragraph association relationship; If the reference extraction coefficient of a text paragraph is less than or equal to the preset reference extraction coefficient, the recommended text of the text paragraph is determined according to the paragraph extraction coefficient.

5. The data recommendation method based on a large model according to claim 4, characterized in that: When segmenting and compensating a text paragraph according to the paragraph association relationship, the content association coefficient of each text paragraph to be analyzed is determined according to the keyword overlap degree and the keyword matching degree, and the retained paragraph of each text paragraph is determined according to the content association coefficient and the text connection coefficient, and the retained paragraph and the recommended text paragraph are recorded as the recommended text of the text paragraph; The recommended text segment is a target text segment whose segment extraction coefficient is greater than a preset segment extraction coefficient, and the to-be-analyzed text segment is a target text segment whose segment extraction coefficient is less than or equal to the preset segment extraction coefficient.

6. The data recommendation method based on a large model according to claim 5, characterized in that: When extracting recommended text content from a search text set by using a combined analysis method, the target search text with the largest search quality index in the search text set is recorded as the key analysis text, and the key keywords are determined according to the application reference values ​​and reference association frequencies of the effective keywords contained in the key analysis text; The target search text containing the key keywords is recorded as the relevant analysis text.

7. The data recommendation method based on a large model according to claim 6, characterized in that: Matching paragraph analysis is performed on each text paragraph in the key analysis text. For a single text paragraph, the matching paragraph analysis process includes: Determine the recommended text paragraphs of the text paragraph and the relevant text paragraphs of each related analysis text according to the key keywords; The matching text segment of the text paragraph is determined according to the matching richness and matching effectiveness of each related text segment.

8. The data recommendation method based on a large model according to claim 7, characterized in that: If the matching text segments of each text paragraph in the key analysis text are determined, the matching segment combination is determined according to the segment overlap coefficient, and the matching key keywords of each matching segment combination are determined according to the distribution reference value and the search relevance; The segment overlap coefficients between the matching text segments in any matching segment combination are all greater than the preset segment overlap coefficient.

Citation Information

Patent Citations

  • Text recommendation method based on machine learning

    CN117195890A

  • Text processing method, related device and equipment

    CN114328852A

  • Text recommendation method and device, equipment and storage medium

    CN116127176A