Intelligent text retrieval analysis method and system that fuses machine language and natural language

By acquiring user search needs and historical data, and using search need tags and historical data represented in machine language, attractiveness evaluation indicators and weights for resources in multiple fields are determined, and matching recommendation values ​​are generated. This solves the problems of insufficient flexibility and accuracy in existing resource retrieval methods and achieves efficient resource recommendation.

CN121597827BActive Publication Date: 2026-04-07GUANGZHOU KEAO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing resource retrieval methods cannot effectively combine machine language and natural language, resulting in insufficient retrieval flexibility and low accuracy, making it difficult to meet complex resource retrieval needs.

Method used

By acquiring user search needs and historical search data, search need tags are determined. Using search need tags represented in machine language and historical search data, comprehensive demand data is determined. Combining attractiveness evaluation indicators and weights of resources in multiple fields, matching recommendation values ​​are generated to recommend resources in multiple fields.

Benefits of technology

It achieves accurate capture of users' true needs using natural language and improves the retrieval efficiency of machine language, avoids retrieval bias caused by ambiguity in natural language, and quickly locks in the range of domain resources that meet the needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597827B_ABST
    Figure CN121597827B_ABST
Patent Text Reader

Abstract

The application relates to the field of text retrieval analysis, and discloses an intelligent text retrieval analysis method and system fusing machine language and natural language. The method comprises the following steps: determining a retrieval demand label according to a user retrieval demand; performing retrieval according to the retrieval demand label and historical retrieval data, and determining comprehensive demand data of multiple field resources; determining multiple attraction evaluation indexes of the multiple field resources according to the retrieval demand label, the historical retrieval data and the comprehensive demand data of the multiple field resources; obtaining multiple evaluation index weights of the field resources; determining the sum of the product of the multiple attraction evaluation indexes of the multiple field resources and the evaluation index weights as matching recommendation values of the multiple field resources; and recommending the multiple field resources in the order from high to low of the matching recommendation values of the multiple field resources. The application avoids retrieval deviation caused by natural language ambiguity through fusion processing of natural language and machine language, and quickly locks the range of field resources meeting the demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text retrieval and analysis technology, and more specifically, to an intelligent text retrieval and analysis method and system that integrates machine language and natural language. Background Technology

[0002] In the field of resource retrieval, existing resource retrieval methods mainly fall into two categories: machine language retrieval and natural language retrieval. Machine language retrieval typically requires the user to input search commands according to preset syntax rules, keyword formats, and other specifications. The retrieval system then parses these machine language commands to accurately match target resources. For example, in a database retrieval scenario, the user needs to input machine language statements containing specific field identifiers and logical operators to limit the search scope. Natural language retrieval, on the other hand, allows the user to input search requests in the form of natural language used in everyday communication, without needing to follow specific format specifications. The retrieval system extracts core search elements through semantic analysis of the natural language. For example, in a search engine retrieval scenario, a user can initiate a search by directly inputting natural language statements such as "how to search for patent documents."

[0003] While existing resource retrieval methods each possess their own advantages, they struggle to effectively integrate machine language and natural language in real-world retrieval scenarios. Current solutions often operate machine language and natural language retrieval independently, lacking an effective collaborative mechanism: machine language retrieval fails to leverage the semantic understanding capabilities of natural language to expand the search scope, resulting in insufficient flexibility; conversely, natural language retrieval struggles to utilize the standardized logic of machine language to improve accuracy. Existing retrieval systems fail to fully integrate the advantages of both methods, hindering the development of efficient retrieval strategies for complex resource searches. This impacts retrieval efficiency and the comprehensiveness of results, ultimately failing to meet users' demands for diversified and efficient resource retrieval. Summary of the Invention

[0004] The purpose of this application is to provide an intelligent text retrieval and analysis method and system that integrates machine language and natural language, solving the technical problem of not being able to perform accurate retrieval by combining machine language and natural language, and achieving the technical effect of being able to perform accurate retrieval by combining machine language and natural language.

[0005] In a first aspect, embodiments of this application provide an intelligent text retrieval and analysis method that integrates machine language and natural language. The method includes: acquiring user retrieval needs and historical retrieval data; determining retrieval need tags based on user retrieval needs; performing a retrieval based on the retrieval need tags and historical retrieval data to determine comprehensive demand data for resources in multiple domains; wherein user retrieval needs are represented in natural language, and retrieval need tags are represented in machine language; determining multiple attractiveness evaluation indicators for resources in multiple domains based on the retrieval need tags, historical retrieval data, and comprehensive demand data for resources in multiple domains; acquiring the weights of multiple evaluation indicators for each domain resource; determining the sum of the products of the multiple attractiveness evaluation indicators and the weights of the evaluation indicators for multiple domain resources as the matching recommendation value for multiple domain resources; and recommending multiple domain resources to the user in descending order of the matching recommendation value for multiple domain resources.

[0006] In one possible implementation, the comprehensive requirement data is determined by searching based on search demand tags and historical search data, including: determining search preference tags based on user search needs and historical search data; obtaining a disambiguation threshold; disambiguating the search demand tags and search preference tags using a rule parsing disambiguation unit configured with the disambiguation threshold to obtain preliminary search instructions; wherein, the search preference tags include high-frequency preference tags and fuzzy query preference tags, and the search preference tags and preliminary search instructions are represented in machine language; obtaining the instruction field weights corresponding to the search preference tags; adjusting the preliminary search instructions based on the instruction field weights to determine structured search instructions; and searching using the structured search instructions to determine comprehensive requirement data for resources in multiple domains; wherein, the structured search instructions are represented in machine language.

[0007] In another possible implementation, a structured search is performed to determine comprehensive demand data for resources across multiple domains. This includes: obtaining cross-language resource keywords and resource summaries for multiple cross-language resources; determining the resource structure features and semantic features of multiple cross-language resources based on the keywords and summaries; performing structure matching calculations on the resource structure features of the cross-language resources according to the structured search instructions to determine the structure matching score of multiple cross-language resources; performing semantic matching calculations on the semantic features of the cross-language resources according to the structured search instructions to determine the semantic matching score of multiple cross-language resources; obtaining the structure matching weight and semantic matching weight; determining the sum of the product of the structure matching score and the structure matching weight, and the product of the semantic matching score and the semantic matching weight, as the comprehensive resource score for multiple cross-language resources; and recommending multiple cross-language resources to the user in descending order of the comprehensive resource score.

[0008] In another possible implementation, the method further includes: acquiring resource domain classifications, resource reference relationships, and domain cooperation data for multiple cross-language resources; performing semantic alignment between user search requirements and multiple cross-language resources to determine search intent and semantic alignment tags; constructing a structure for the semantic alignment tags and resource domain classifications to determine a cross-domain search index for multiple cross-language resources; performing matching calculations between the search intent and the cross-domain search index to obtain content matching scores for multiple cross-language resources; performing association quantification calculations between the search intent and resource reference relationships and domain cooperation data for multiple cross-language resources to obtain domain association scores for multiple cross-language resources; acquiring content matching weights and domain association weights; determining the sum of the product of the content matching scores and content matching weights, and the product of the domain association scores and domain association weights, as the comprehensive resource score for multiple cross-language resources; and recommending multiple cross-language resources to the user in descending order of the comprehensive resource scores.

[0009] In another possible implementation, a structure is constructed for semantic alignment labels and resource domain classifications to determine cross-domain retrieval indexes for multiple cross-language resources. This includes: extracting features from professional documents and structured data to obtain professional element features; where professional element features include professional clause numbers, data fields, and core keywords; constructing a graph of professional element features to obtain a professional knowledge graph; and constructing a unified index for the professional knowledge graph to obtain a cross-domain retrieval index.

[0010] In another possible implementation, the method further includes: obtaining historical search risk cases where resource search results and user search needs do not match; extracting risk dimensions from historical search risk cases to determine the domain risk dimension features of historical search risk cases; defining rules for the domain risk dimension features to determine risk dimension weight rules; wherein, historical search risk cases include professional document search results, structured data search results, and user search needs; determining risk dimension scores for multiple domain resources through risk dimension weight rules; obtaining matching recommendation weights and risk dimension weights; determining the difference between the product of the matching recommendation value and the matching recommendation weight of multiple domain resources, and the product of the risk dimension score and the risk dimension weight, as the corrected matching recommendation value for multiple domain resources; and recommending multiple domain resources to the user in descending order of the corrected matching recommendation value of multiple domain resources.

[0011] In another possible implementation, the method further includes: determining the sum of the products of structural matching scores and structural matching weights, and semantic matching scores and semantic matching weights of multiple cross-language resources, as a comprehensive resource score for the multiple cross-language resources; determining the difference between the products of the comprehensive resource score, risk dimension score, and risk dimension weight of the multiple cross-language resources, as a revised comprehensive resource score for the multiple cross-language resources; and recommending multiple cross-language resources to the user in descending order of the revised comprehensive resource scores.

[0012] In another possible implementation, the method further includes: determining the sum of the products of content matching scores and content matching weights, and the products of domain association scores and domain association weights of multiple cross-language resources, as a comprehensive resource score for the multiple cross-language resources; determining the difference between the products of the comprehensive resource score, risk dimension score, and risk dimension weight of the multiple cross-language resources, as a corrected comprehensive resource score for the multiple cross-language resources; and recommending multiple cross-language resources to the user in descending order of the corrected comprehensive resource scores.

[0013] In another possible implementation, based on search demand tags, historical search data, and comprehensive demand data from multiple domains, several attractiveness evaluation indicators for resources in multiple domains are determined. These include: determining the similarity between search demand tags and tag text of resources in multiple domains, and determining the historical search similarity between historical search data and historical search data of resources in multiple domains; obtaining resource timeliness score, resource authority score, resource completeness score, and user preference score for resources in multiple domains; wherein, the user preference score is used to characterize the score corresponding to the click-through conversion rate and average dwell time ratio of domain resources; and using the tag text similarity, historical search similarity, resource timeliness score, resource authority score, resource completeness score, and user preference score of resources in multiple domains as multiple attractiveness evaluation indicators for resources in multiple domains. Obtain the weights of multiple evaluation indicators for resources in various fields; determine the sum of the products of multiple attractiveness evaluation indicators and their weights for resources in multiple fields, as the matching recommendation value for resources in multiple fields, including: obtaining the tag text similarity weight, historical search similarity weight, resource timeliness weight, resource authority weight, resource completeness weight, and user preference weight for resources in each field; determine the sum of the products of tag text similarity and tag text similarity weight, historical search similarity and historical search similarity weight, resource timeliness score and resource timeliness weight, resource authority score and resource authority weight, resource completeness score and resource completeness weight, and user preference score and user preference weight for resources in multiple fields, as the matching recommendation value for resources in multiple fields.

[0014] Secondly, embodiments of this application provide an intelligent text retrieval and analysis system that integrates machine language and natural language, including units for implementing the above-described method.

[0015] The beneficial effects of the embodiments in this application compared with the prior art are:

[0016] This application provides an intelligent text retrieval and analysis method that integrates machine language and natural language. The method includes: acquiring user search needs and historical search data; determining search need tags based on user search needs; performing searches based on search need tags and historical search data to determine comprehensive demand data for resources in multiple domains; determining multiple attractiveness evaluation indicators for resources in multiple domains based on search need tags, historical search data, and comprehensive demand data for resources in multiple domains; acquiring the weights of multiple evaluation indicators for each domain resource; determining the product of the attractiveness evaluation indicators and the weights of the evaluation indicators for multiple domain resources as the matching recommendation value for multiple domain resources; and recommending multiple domain resources to the user in descending order of the matching recommendation values ​​for multiple domain resources. In this application embodiment, by integrating natural language and machine language, it can accurately capture the user's true search needs, improve search efficiency by leveraging the standardization of machine language, avoid search biases caused by ambiguity in natural language, and quickly pinpoint the range of domain resources that meet the user's needs. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating the first intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0019] Figure 2 A schematic diagram illustrating the workflow of the first intelligent text retrieval and analysis method integrating machine language and natural language provided in the embodiments of this application;

[0020] Figure 3 A flowchart illustrating the second intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0021] Figure 4 A schematic diagram illustrating the workflow of the second intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0022] Figure 5 A flowchart illustrating the third intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0023] Figure 6 A schematic diagram illustrating the workflow of the third intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0024] Figure 7 A flowchart illustrating the fourth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0025] Figure 8 A schematic diagram illustrating the workflow of the fourth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0026] Figure 9 A flowchart illustrating the fifth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0027] Figure 10 A flowchart illustrating the sixth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0028] Figure 11 A schematic diagram illustrating the workflow of the sixth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0029] Figure 12 A flowchart illustrating the seventh intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0030] Figure 13 A flowchart illustrating the eighth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment;

[0031] Figure 14 This is a schematic diagram of the logical structure of an intelligent text retrieval and analysis system that integrates machine language and natural language, provided as an embodiment of this application. Detailed Implementation

[0032] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0033] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0034] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0035] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0036] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0037] Existing retrieval systems cannot fully integrate the advantages of the two retrieval methods. When faced with complex resource retrieval needs, they are unable to form efficient retrieval strategies, which in turn affects retrieval efficiency and the comprehensiveness of retrieval results.

[0038] Based on the above reasons, this application provides an intelligent text retrieval and analysis method that integrates machine language and natural language. The method includes: acquiring user search needs and historical search data; determining search need tags based on user search needs; performing searches based on search need tags and historical search data to determine comprehensive demand data for resources in multiple domains; determining multiple attractiveness evaluation indicators for resources in multiple domains based on search need tags, historical search data, and comprehensive demand data for resources in multiple domains; acquiring the weights of multiple evaluation indicators for each domain resource; determining the product of the attractiveness evaluation indicators and the weights of the evaluation indicators for multiple domain resources as the matching recommendation value for multiple domain resources; and recommending multiple domain resources to the user in descending order of the matching recommendation values ​​for multiple domain resources. In this application embodiment, by integrating natural language and machine language, it can accurately capture the user's true search needs, improve search efficiency by leveraging the standardization of machine language, avoid search biases caused by ambiguity in natural language, and quickly lock in the range of domain resources that meet the user's needs.

[0039] In some scenarios, the intelligent text retrieval and analysis method that integrates machine language and natural language according to the embodiments of this application can be applied to resource retrieval in the fields of scientific research, medicine, and law, and can improve the accuracy of resource retrieval in these fields.

[0040] The following describes in detail, with specific examples, an intelligent text retrieval and analysis method that integrates machine language and natural language, provided in the embodiments of this application.

[0041] Figure 1 A flowchart illustrating the intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 1 As shown in the figure, this application provides an intelligent text retrieval and analysis method that integrates machine language and natural language. The method includes steps S110 to S120, which will be described in detail below.

[0042] S110. Obtain user search requirements and historical search data. Based on user search requirements, determine search requirement tags. Perform searches based on search requirement tags and historical search data to determine comprehensive demand data for resources across multiple domains. User search requirements are represented in natural language, while search requirement tags are represented in machine language.

[0043] Figure 2 A schematic diagram illustrating the workflow of the first intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 2As shown, in this implementation, the search requests submitted by users can be collected through a preset interactive interface, and the historical search records of the user or similar users stored in the system can be retrieved and integrated to form a basic dataset for subsequent search analysis.

[0044] For example, user search requests can be collected through web input boxes or mobile voice interaction modules, and historical search data such as the user's search keywords, browsing time, and resource download records over a period of time can be extracted from the system database.

[0045] In this implementation, natural language processing technology can be used to perform semantic parsing, entity recognition, and intent extraction on user search requests, converting unstructured request content into standardized search request tags. This not only accurately captures the user's true search needs but also improves search efficiency by leveraging the standardization of machine language, avoiding search biases caused by ambiguity in natural language.

[0046] For example, the natural language requirements submitted by users can be segmented to extract information such as professional terms, industry categories, and topic directions corresponding to the core requirements, and mapped into a set of standardized tags that can be recognized by machines.

[0047] It should be noted that user search requests are expressed in natural language, while search request tags are expressed in machine language. Users can express their search needs using everyday spoken or written natural language, without needing to follow a specific format; search request tags, on the other hand, adopt machine-standard structured coding or standardized terminology to facilitate efficient system processing and matching.

[0048] For example, a user can input a natural language retrieval request for learning materials in a specific field, and the system will convert it into a set of machine language tags containing the corresponding field, learning stage, and material type.

[0049] For example, a user's search request in natural language representation could be: core journal articles related to simulation modeling of lithium-ion battery thermal management systems for new energy vehicles. A search request tag in machine language representation could be: {Subject Classification Code: 080502; Research Direction: New Energy Vehicles | Lithium-ion Batteries | Thermal Management Systems | Simulation Modeling; Document Type: Core Journals; Publication Time: Last 5 Years}.

[0050] In this implementation, search requirement tags can be used as core search conditions. Combined with historical search data, domain resources with high relevance to user needs can be selected. Then, the relevant attribute data of these resources can be integrated to form a comprehensive requirement dataset, which can quickly lock in the range of domain resources that meet the requirements.

[0051] For example, resources in the corresponding field can be located in the knowledge base based on search demand tags. By combining the resource types preferred by users and update time in historical search data, target resources can be filtered out and their attribute data such as topic classification, content depth, and user feedback can be extracted.

[0052] S120. Based on search demand tags, historical search data, and comprehensive demand data from multiple domains, determine multiple attractiveness evaluation indicators for resources in multiple domains. Obtain the weights of the multiple evaluation indicators for each domain resource. Determine the product of the attractiveness evaluation indicators and the weights of the evaluation indicators for multiple domain resources, and use this as the matching recommendation value for multiple domain resources. Recommend multiple domain resources to the user in descending order of their matching recommendation values.

[0053] like Figure 2 As shown, in this implementation method, a multi-dimensional attractiveness evaluation index system can be built based on the core demands extracted from search demand tags, user preferences reflected in historical search data, and resource attributes contained in comprehensive demand data, covering aspects such as the fit between resources and demands and potential user preferences.

[0054] For example, resource theme matching degree, content depth adaptation degree, user recognition degree, etc. can be set as attractiveness evaluation indicators, and each indicator corresponds to specific quantitative judgment rules.

[0055] For example, when determining multiple attractiveness evaluation indicators for resources in multiple fields based on search demand tags, historical search data, and comprehensive demand data for resources in multiple fields, multiple attractiveness evaluation indicators for resources in multiple fields can be determined by using an empirical value table corresponding to search demand tags, historical search data, comprehensive demand data for resources in multiple fields, and multiple attractiveness evaluation indicators for resources in multiple fields.

[0056] In this implementation, the weight values ​​corresponding to each attractiveness evaluation index can be determined based on the coreness of the search needs and the user priority reflected in historical search data through methods such as the analytic hierarchy process and user preference learning models, so as to provide a quantitative basis for subsequent matching recommendation value calculation.

[0057] For example, if historical search data reveals that users are more concerned about the content depth of resources, the weight of the content depth fit index is increased, while the weight of other secondary indexes is decreased.

[0058] In this implementation, the attractiveness evaluation indicators of each resource in each field can be multiplied with their corresponding weights, and then all the product results can be summed to obtain the matching recommendation value of the resource, thereby objectively measuring the degree of fit between resources in each field and user needs.

[0059] For example, the matching recommendation value of a resource is obtained by multiplying the topic matching degree index value and weight of a resource in a certain field with the content depth adaptation degree index value and weight, and then summing the results.

[0060] In this implementation, the matching recommendation values ​​of all domain resources can be sorted in descending order, and the corresponding domain resources can be displayed to the user in sequence according to the sorting results. This ensures that the priority of the recommended resources meets the user's core needs and improves the rationality and relevance of the recommendation results.

[0061] For example, after sorting the matching recommendation values ​​from high to low, the top-ranked resources in several fields are displayed first, and the core matching points of each resource are marked to help users quickly understand the advantages of the resources.

[0062] This implementation converts user search requests represented in natural language into search request tags represented in machine language. By combining search request tags with historical search data, it determines the comprehensive demand data for resources in multiple fields. This processing method, which integrates natural language and machine language, can accurately capture users' real search needs, improve search efficiency by leveraging the standardization of machine language, avoid search biases caused by ambiguity in natural language, and quickly lock in the range of field resources that meet the needs.

[0063] This implementation method, through the quantitative calculation of multi-dimensional evaluation indicators and corresponding weights, can objectively measure the degree of fit between resources and user needs in various fields, ensuring that the priority of recommended resources meets the core needs of users and improving the rationality and relevance of recommendation results. The full-process multi-data linkage mechanism can realize closed-loop optimization from demand capture to resource recommendation, effectively integrating the convenience of natural language interaction and the accuracy of machine language processing, providing users with domain resource recommendation services that are more in line with their own needs.

[0064] Figure 3 A flowchart illustrating the second intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 3 As shown, in some implementations, in the above-mentioned S110, the comprehensive requirement data is determined based on the search requirement tags and historical search data, including S111 to S112. S111 to S112 will be explained in detail below.

[0065] S111. Based on user search needs and historical search data, determine search preference tags. Obtain the disambiguation threshold. Using the rule parsing disambiguation unit configured with the disambiguation threshold, disambiguate the search need tags and search preference tags to obtain preliminary search instructions. The search preference tags include high-frequency preference tags and fuzzy query preference tags. The search preference tags and preliminary search instructions are represented in machine language.

[0066] Figure 4 A schematic diagram illustrating the workflow of the second intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 4 As shown, in this implementation, search preference tags can be determined by combining user search needs and historical search data.

[0067] For example, frequently occurring search terms and related terms when users make broad queries can be filtered from users' historical search data to generate corresponding search preference tags.

[0068] In this implementation, the disambiguation threshold for label disambiguation processing can be obtained through a preset system configuration module.

[0069] For example, a reasonable disambiguation threshold value can be set based on the historical disambiguation effect of the retrieval system to provide a basis for judgment in subsequent label disambiguation operations.

[0070] In this implementation, the rule parsing disambiguation unit configured with the disambiguation threshold can be used to disambiguate the search request tags and search preference tags to obtain preliminary search instructions. Search preference tags can include high-frequency preference tags and fuzzy query preference tags. Both the search preference tags and the preliminary search instructions are represented using machine language.

[0071] It should be noted that this disambiguation process can effectively resolve ambiguity or conflicting information between search tags, filter invalid tag content, generate preliminary search instructions that are more in line with the user's actual needs, and reduce the probability of errors in subsequent searches.

[0072] For example, vague descriptions in search query tags can be corrected by combining high-frequency preference tags, and tag content that does not conform to the user's long-term search habits can be eliminated to obtain accurate preliminary search instructions.

[0073] For example, a semantic similarity threshold of ≥85% can be set as the criterion for determining valid tags; and a historical search frequency ratio of ≥30% can be set as the criterion for retaining high-frequency preference tags. If the semantic similarity between the search request tag and the search preference tag is ≥85%, they are determined to be homologous tags, retained, and merged; if it is <85%, they are determined to be ambiguous tags and removed.

[0074] In this implementation, high-frequency preference tags must meet the requirement of "historical search frequency accounting for ≥30%", otherwise they are marked as low-value preference tags, and their priority is reduced or they are removed.

[0075] In this implementation, the matching between the fuzzy query preference tags and the core requirement tags needs to be verified. If the fuzzy rule may cause the search scope to deviate from the core requirement (e.g., suffix fuzzy matching may include irrelevant content such as "battery material pollution"), then the fuzzy rule will be removed.

[0076] For example, the initial search command could be { "domain": "new energy vehicles", "object": ["battery materials"], "target": ["lifespan optimization"], "preference_high_freq": ["ternary lithium batteries", "lithium iron phosphate"], "preference_fuzzy": ["prefix_match"]}.

[0077] S112. Obtain the command field weights corresponding to the search preference tags. Adjust the initial search commands based on the command field weights to determine the structured search commands. Perform a search using the structured search commands to determine the comprehensive demand data for resources across multiple domains. The structured search commands are represented in machine language.

[0078] In this implementation, the weight of the corresponding instruction field can be obtained based on the type of search preference tag and the weight ratio in the user's search history.

[0079] For example, a higher instruction field weight can be configured for high-frequency preference tags, and an appropriate weight value can be configured for fuzzy query preference tags to reflect the priority differences of different tags.

[0080] In this implementation, the obtained instruction field weights can be used to adjust the priority of each tag field in the initial search instruction, ultimately determining the structured search instruction. The structured search instruction is represented using machine language.

[0081] It should be noted that by configuring exclusive command field weights for different search preference tags, the content of tags that users frequently pay attention to can be strengthened in a targeted manner, while the influence of secondary tags can be weakened. This makes the structured search commands more in line with users' personalized search preferences, thereby improving the accuracy and targeting of search commands.

[0082] For example, the weight of the instruction field corresponding to high-frequency preference tags can be applied to the initial search instruction to amplify the influence of such tags in the search and generate a structured search instruction that meets the user's personalized needs.

[0083] For example, a structured retrieval instruction could be: { "instruction_type": "structured_retrieval", "fields": [ {"name": "domain", "value": "new energy vehicles", "weight":1.0, "match_type": "exact"}, {"name": "object", "value": ["battery materials", "ternary lithium batteries", "lithium iron phosphate"], "weight": [1.0, 1.2, 1.0], "match_type": "exact"}, {"name": "target", "value": "lifespan optimization", "weight": 1.0, "match_type": "exact"}], "global_rule": { "match_mode": "prefix_fuzzy", "rule_weight": 1.5, "apply_scope": ["object", "target"]}, "language_type": "machine_readable"}. Among them, the high-frequency preference tag is ternary lithium battery, with the field "object" and a weight of 1.2, which is higher than the weight of the basic tag and is matched first. The high-frequency preference tag is lithium iron phosphate, with the field "object" and a weight of 1.0, which is the same as the weight of the basic tag. The fuzzy query preference tag is prefix fuzzy matching, with the field "match_mode" and a weight of 1.5, which forcibly increases the priority of fuzzy matching. During the search process, high-weight fields are matched first, and resources containing "ternary lithium battery" and associated with "lifespan optimization" under the "new energy vehicle" field are searched first (such as patent documents, industry reports, experimental data). By applying fuzzy matching rules, prefix fuzzy matching is performed on the "object" and "target" fields, such as matching derivative keywords such as "battery materials" → "solid-state battery materials" and "lifespan optimization" → "lifespan improvement".

[0084] In this implementation, the generated structured search instructions can be used to search in a preset domain resource database to determine the comprehensive demand data for resources in multiple domains.

[0085] It should be noted that the structured search command has undergone dual optimization of disambiguation and weight adjustment. Compared with the method of directly using search demand tags and historical search data, it can more accurately match the user's real needs. The comprehensive demand data obtained by the search has a higher degree of consistency with the user's demands, providing more reliable data support for the subsequent determination of attractiveness evaluation indicators and resource recommendations.

[0086] For example, a structured search command can be input into the search system, and the system can filter out domain resources that meet the requirements of various tags in the database based on the command, and summarize the corresponding comprehensive demand data.

[0087] This implementation determines search preference tags based on user search needs and historical search data, obtains a disambiguation threshold, and uses a rule-based disambiguation unit configured with that threshold to disambiguate the search need tags and search preference tags, generating a preliminary search instruction. The search preference tags include high-frequency preference tags and fuzzy query preference tags, and both the search preference tags and the preliminary search instruction are represented in machine language. This disambiguation process effectively resolves ambiguity or conflicting information between search tags, filters invalid tag content, generates a preliminary search instruction that better reflects the user's actual needs, and reduces the probability of errors in subsequent searches.

[0088] This implementation obtains the weights of instruction fields corresponding to search preference tags, adjusts the initial search instructions based on these weights, and determines the structured search instructions, which are then represented in machine language. By configuring dedicated instruction field weights for different search preference tags, it is possible to specifically strengthen tags that users frequently focus on and weaken the influence of secondary tags, making the structured search instructions more aligned with users' personalized search preferences and improving the accuracy and targeting of the search instructions.

[0089] This implementation method uses structured search commands to determine comprehensive demand data for resources across multiple fields. The structured search commands have undergone dual optimization through disambiguation and weight adjustment. Compared to directly using search demand tags and historical search data, this method can more accurately match users' actual needs. The comprehensive demand data obtained from the search has a higher degree of consistency with user requests, providing more reliable data support for subsequent determination of attractiveness evaluation indicators and resource recommendations.

[0090] Figure 5 A flowchart illustrating the third intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 5 As shown, in some implementations, in the above S112, a structured search instruction is used to search and determine the comprehensive demand data of resources in multiple fields, including S112a to S112b. S112a to S112b will be explained in detail below.

[0091] S112a. Obtain cross-language resource keywords and resource summaries from multiple cross-language resources. Based on the cross-language resource keywords and resource summaries, determine the resource structure features and semantic features of the multiple cross-language resources.

[0092] Figure 6A schematic diagram illustrating the workflow of the third intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 6 As shown, this implementation method can collect core keywords of various resources in multiple languages ​​in batches, as well as resource summaries that extract the core content of the resources, which can provide a data foundation for accurately matching user needs.

[0093] In this implementation, the collected cross-language resource keywords and resource summaries can be structured and semantically mined, for example, extracting the structural features of the resources from the classification hierarchy and arrangement rules of the keywords, and extracting the semantic features of the resources from the theme expression and core appeal of the summaries.

[0094] It should be noted that by extracting structural and semantic features from keywords and summaries of cross-linguistic resources, we can comprehensively explore the deep attributes of the resources, break through the limitations of single-dimensional evaluation, and provide multi-dimensional support for subsequent accurate matching.

[0095] In this implementation, cross-language resources can include three dimensions: resource language, resource keywords, and resource abstracts. Among them, Chinese resource keywords focus on new energy vehicles and ternary lithium batteries, and the abstracts revolve around the optimization strategies and test results of ternary lithium battery cycle life. English resource keywords involve lithium iron phosphate batteries and temperature control, and the abstracts describe the research on improving the lifespan of lithium iron phosphate batteries through temperature control. Japanese resource keywords include lead-acid batteries and charging control, and the abstracts focus on the research on charging control of lead-acid batteries for electric vehicles.

[0096] For example, for a certain English technical literature resource, structural features can be extracted from the classification tag level of its keywords and the logical association between chapters, and semantic features can be extracted from the core objectives and application scenarios of the technical solutions described in the abstract.

[0097] For example, in the extraction of resource structure features in machine language JSON format, the extraction rules can be divided into four dimensions: domain-level keywords, object-level keywords, target-level keywords, and summary structure type, marking the existence and quantity of each type of element.

[0098] For example, in the extraction of semantic features of resources in machine language JSON format, the extraction rule can be to use a cross-language semantic alignment model (such as LASER) to map non-Chinese keywords / summaries to the Chinese semantic space and extract core concepts and concept associations.

[0099] S112b. Based on the structured search instructions, perform structure matching calculations on the resource structure features of cross-language resources to determine the structure matching scores of multiple cross-language resources. Based on the structured search instructions, perform semantic matching calculations on the semantic features of cross-language resources to determine the semantic matching scores of multiple cross-language resources. Obtain the structure matching weights and semantic matching weights. Determine the sum of the products of the structure matching scores and structure matching weights, and the semantic matching scores and semantic matching weights, as the comprehensive resource score for multiple cross-language resources. Recommend multiple cross-language resources to the user in descending order of their comprehensive resource scores.

[0100] In this implementation, the generated structured search instructions can be used as a matching benchmark and compared with the resource structure features of each cross-language resource in the corresponding dimension. Through a standardized structure matching algorithm, the fit between the resource structure and the search instruction structure requirements is quantitatively evaluated, and a structure matching score for the corresponding cross-language resource is generated.

[0101] For example, the structure matching score calculation process can be carried out around three cross-language resources: Chinese, English, and Japanese. Scores are accumulated based on four matching rules and their corresponding weights: "existence of domain-related keywords, number of intersections between object-related keywords and the search command list, existence of target-related keywords, and completeness of the summary structure." For instance, the example Chinese and example English resources both satisfy the conditions of having domain-related keywords (2 points each), having an intersection between object-related keywords and the search command object list (4 points each), having target-related keywords (2 points each), and having a complete summary structure (2 points each). The final structure matching score for both is 10 points. The example Japanese resource, while satisfying the conditions of having domain-related and target-related keywords (2 points each), has no intersection between object-related keywords and the command (0 points) and an incomplete summary structure (1 point). The final structure matching score is 5 points.

[0102] In this implementation, the semantic requirements carried by the structured search instructions can be used as a reference to calculate the similarity with the semantic features of various cross-language resources. Through semantic analysis algorithms, the degree of matching between the semantics of the resources and the core requirements of the search instructions can be measured to obtain the semantic matching score of the corresponding cross-language resources.

[0103] For example, the semantic matching score calculation process can be based on a cross-language semantic similarity algorithm. First, the core semantics of the structured search instruction (new energy vehicles, ternary lithium batteries / lithium iron phosphate, lifespan optimization) are constructed as semantic vectors. Then, the cosine similarity between the core concept vectors of the three types of resources and the semantic vector of the structured search instruction is calculated separately. Finally, the semantic matching score is obtained through a conversion method of "similarity value × 10". For example, the cosine similarity between the core concepts of Chinese resources and the core semantics of the structured search instruction is the highest, reaching 0.95, corresponding to a semantic matching score of 9.5. The similarity between English resources is 0.90, resulting in a converted score of 9.0. The core concepts of Japanese resources and the core semantics of the structured search instruction have a lower similarity, only 0.40, resulting in a final semantic matching score of 4.0.

[0104] It should be noted that by splitting the matching process into two independent dimensions, structural matching and semantic matching, the weight ratio can be flexibly adjusted according to the needs of different retrieval scenarios, so as to meet diverse retrieval requirements and improve the flexibility and personalization of resource matching.

[0105] For example, in academic literature retrieval scenarios, if users pay more attention to the structural standardization of documents, the weight of structural matching can be increased; if users pay more attention to the thematic relevance of documents, the weight of semantic matching can be increased.

[0106] In this implementation, the structural matching weight and semantic matching weight adapted to the current retrieval scenario can be retrieved. These weight parameters can be preset or dynamically adjusted according to the needs of different retrieval scenarios to balance the influence of structural and semantic dimensions in the comprehensive evaluation.

[0107] In this implementation, a weighted fusion calculation can be performed on each cross-language resource. First, its structural matching score and structural matching weight are multiplied together. Then, its semantic matching score and semantic matching weight are multiplied together. The sum of the two products is used as the overall resource score for that cross-language resource. This score can comprehensively reflect the overall fit between the resource and the search requirements.

[0108] In this implementation, the comprehensive resource scores of all cross-language resources can be sorted in descending order to form an ordered resource recommendation sequence. The sorted resource sequence is then pushed to the user interface, allowing users to prioritize cross-language resources that best match their search needs.

[0109] This implementation first obtains cross-language resource keywords and resource summaries from multiple cross-language resources. Based on this, it determines the resource structure and semantic features of multiple cross-language resources. Then, based on structured search commands, it performs structure matching calculations on the resource structure features to obtain a structure matching score, and performs semantic matching calculations on the semantic features to obtain a semantic matching score. Finally, it combines the structure matching weight and the semantic matching weight, and obtains a comprehensive resource score through weighted summation. This approach can comprehensively uncover the deep attributes of cross-language resources, making the matching between cross-language resources and structured search commands more accurate and improving the accuracy of cross-language resource retrieval.

[0110] This implementation breaks down matching into two independent dimensions: structure and semantics. The weighting ratio can be flexibly adjusted according to different search scenarios to meet diverse search needs and improve the flexibility and personalization of resource matching. By comprehensively considering the structural and semantic adaptability of resources, it can prioritize the push of cross-language resources that best meet the user's search needs, shorten the time for users to obtain effective resources, and improve the efficiency of search recommendations and user experience.

[0111] Figure 7 A flowchart illustrating the fourth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 7 As shown, in some implementations, the above method also includes S210 to S230, which will be described in detail below.

[0112] S210. Obtain resource domain classification, resource reference relationships, and domain cooperation data for multiple cross-language resources; perform semantic alignment between user search requirements and multiple cross-language resources to determine search intent and semantic alignment tags; construct a structure for semantic alignment tags and resource domain classification to determine the cross-domain search index for multiple cross-language resources.

[0113] Figure 8 A schematic diagram illustrating the workflow of the fourth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 8 As shown, structurally, functionally, and in use,

[0114] In this implementation, resource domain classification, resource citation relationships, and domain cooperation data of multiple cross-language resources can be obtained from metadata of cross-language resources, academic database association information, etc., providing multi-dimensional basic data support for subsequent retrieval matching and scoring calculation.

[0115] It should be noted that resource domain classification can clearly identify the sub-domain to which cross-language resources belong, resource citation relationships can reflect the industry recognition of resources, and domain cooperation data can reflect the collaboration between institutions behind the resources. These three types of data together constitute a multi-dimensional attribute profile of cross-language resources.

[0116] For example, information such as the domain classification tags, citation counts, and co-publishing institutions of foreign language literature can be obtained from a global academic journal database as attribute data for corresponding cross-language resources.

[0117] In this implementation, a multilingual semantic alignment model can be used to perform semantic matching and calibration of user search requirements and multiple cross-language resources, determine the accurate search intent and corresponding semantic alignment tags, and resolve semantic differences between languages.

[0118] In this implementation, the search intent refers to the core information needs extracted from the user's search request statement. It is the goal orientation of the user's search behavior and must clearly define key elements such as the scope of the search field, technical direction, and information type. It is not affected by differences in language expression.

[0119] It should be noted that the semantic alignment process ensures that the semantic understanding of user search needs and cross-language resources remains consistent, avoiding deviations in subsequent matching calculations due to differences in the expression habits of different languages.

[0120] For example, for a user's search request in the field of computer vision submitted in Chinese, it can be semantically aligned with cross-language documents in English and German to determine a consistent search intent and semantic alignment tags.

[0121] For example, a multilingual semantic alignment model can be an empirical model based on the summary of empirical value tables. Through the multilingual semantic alignment model, the user's search needs and multiple cross-language resources can be semantically aligned to determine the search intent and semantic alignment tags.

[0122] For example, if a user's search query includes the Chinese phrase "cross-border e-commerce supply chain resilience construction strategy" and the English phrase "Resilience construction strategies for cross-border e-commerce supply chains", a multilingual semantic parsing model can be used to separate the differences in language expressions, extract the core requirements, and clarify the specific strategies, practical cases, and related assessment methods for cross-border e-commerce supply chain resilience construction. This is the search intent.

[0123] In this implementation, semantic alignment tags serve as a standardized medium connecting user search intent with cross-language resource features. By semantically mapping the domain classification, technical features, and content themes of cross-language resources, a set of tags matching the search intent is generated, eliminating semantic and expressive differences between languages, and achieving precise association between search intent and cross-language resources.

[0124] For example, thematic semantic extraction can be performed on cross-language resources such as Chinese journal articles and English industry reports. The core concepts in user needs, such as "cross-border e-commerce," "supply chain resilience," and "building strategies," can be semantically mapped with multilingual expressions in resources, such as "Cross-border E-commerce," "Supply Chain Resilience," and "Strategy Formulation." This generates semantic alignment tags covering the domain layer (cross-border e-commerce, supply chain management), the content layer (resilience building strategies, risk response cases), and the language mapping layer (CN - cross-border e-commerce / EN - Cross-border E-commerce), thus completing the semantic connection between needs and resources.

[0125] In this implementation, semantic alignment tags and resource domain classifications can be structurally integrated according to preset index structure rules to determine cross-domain retrieval indexes for multiple cross-language resources, providing a standardized retrieval basis for subsequent content matching calculations.

[0126] S220. Match the search intent with the cross-domain search index to obtain content matching scores for multiple cross-language resources. Perform correlation quantification calculations on the search intent and resource reference relationships and domain cooperation data of multiple cross-language resources to obtain domain association scores for multiple cross-language resources.

[0127] In this implementation, a preset content matching algorithm can be used to perform quantitative matching calculations on the search intent and cross-domain search index, and obtain content matching scores for multiple cross-language resources, which intuitively reflects the degree of fit between resource content and user search needs.

[0128] It should be noted that the content matching score is calculated based on the direct matching degree between the cross-domain search index and the search intent. The higher the score, the more closely the resource content matches the user's core search needs.

[0129] For example, when the search intent focuses on deep learning model optimization, the intent can be matched with cross-domain search indexes of cross-language resources to calculate the corresponding content matching score.

[0130] For example, taking cross-language resource retrieval for AI-driven medical image diagnosis as an example, when determining the search intent and calculating the content matching score by matching the cross-domain search index, the search intent is "to obtain application patents and technical documents of deep learning-driven image recognition algorithms in CT and MRI image lesion diagnosis". The corresponding cross-domain search index is a three-level hierarchical index constructed based on semantic alignment tags and resource domain classification, specifically including domain layer index (AI - Computer Vision / Biomedicine - Medical Image Diagnosis), technology layer index (convolutional neural network algorithm / lesion feature extraction / multimodal image fusion), and application layer index (CT image diagnosis / MRI image diagnosis). A hierarchical weighted matching algorithm was used for matching calculations, with matching weights of 0.4, 0.4, and 0.2 for the domain layer, technology layer, and application layer, respectively. The index level of each cross-language resource was compared one by one with the corresponding level of the search intent. If the cross-domain search index of a Chinese patent resource fully covers all three levels and the technology layer index explicitly states "Application of Convolutional Neural Network in Lung CT Lesion Recognition", then the domain layer matching score of this resource is 0.4, the technology layer matching score is 0.4, the application layer matching score is 0.2, and the total content matching score is 1.0. If the index of an English paper resource only covers the domain layer and the technology layer, and the application layer is labeled as "X-ray Image Diagnosis", which does not match the CT and MRI in the search intent, then its application layer score is 0, and the total content matching score is 0.8.

[0131] In this implementation, a domain association quantification model can be used to perform association analysis and quantification calculation on the search intent and the resource reference relationship and domain cooperation data of multiple cross-language resources, so as to obtain the domain association score of multiple cross-language resources and explore the domain reference value of the resources.

[0132] It should be noted that the domain relevance score can reflect the influence and expansion value of cross-language resources within their respective domains, breaking through the limitations of evaluation based on a single content dimension and providing users with more in-depth resource references.

[0133] For example, when a user searches for resources in the field of new energy materials, the search intent can be correlated with the citation frequency of cross-language resources and cooperation data of companies in the same field to obtain the corresponding field correlation score.

[0134] For example, when determining the search intent and resource citation relationship, and quantifying the domain association score based on domain collaboration data, the search intent implicitly includes the potential need to "prioritize resources with high recognition and strong technological applicability in the medical field." Therefore, the association quantification calculation revolves around the resource citation relationship and domain collaboration data: In the domain association quantification model, the resource citation relationship weight can be set to 0.5, the domain collaboration data weight to 0.5, and the citation relationship score is converted based on the number of citations of high-impact literature in the same field, with 0.2 points for each citation by a top medical imaging journal and 0.1 points for each citation by a general computer vision journal; the domain collaboration data score is converted based on the depth of medical field collaboration with the institution to which the resource belongs, with 0.15 points for each collaboration with a top-tier hospital or authoritative medical research institution. Regarding the aforementioned Chinese patent resources, they were cited in 5 top medical imaging journals, and the affiliated research institution had joint research collaborations with 3 tertiary hospitals. Therefore, the citation relationship score is 5 × 0.2 = 1.0, the field collaboration score is 3 × 0.15 = 0.45, and the field association score is calculated as (1.0 × 0.5) + (0.45 × 0.5) = 0.725. Regarding the aforementioned English paper resources, they were only cited in 2 ordinary computer vision journals, and the affiliated institution had no medical field collaboration records. Therefore, the citation relationship score is 2 × 0.1 = 0.2, the field collaboration score is 0, and the field association score is calculated as (0.2 × 0.5) + (0 × 0.5) = 0.1.

[0135] S230. Obtain content matching weight and domain relevance weight. Determine the sum of the products of content matching scores and content matching weights, and the products of domain relevance scores and domain relevance weights for multiple cross-language resources, as the comprehensive resource score for the multiple cross-language resources. Recommend multiple cross-language resources to the user in descending order of their comprehensive resource scores.

[0136] In this implementation, preset or dynamically adjusted content matching weights and domain association weights can be obtained based on the user's historical search preferences and the domain attributes of the current search scenario, providing a basis for dimensional priority in the comprehensive score calculation.

[0137] It should be noted that the settings for content matching weight and domain relevance weight can be flexibly adjusted according to the user's search needs, balancing the relevance of resource content and the value of domain relevance.

[0138] For example, in scenarios where users need to accurately obtain technical information, a higher content matching weight can be set; in scenarios where users need to explore cutting-edge developments in a field, a higher field relevance weight can be set.

[0139] In this implementation, the product of content matching score and content matching weight, and the product of domain association score and domain association weight can be calculated separately. The two product results are then summed to obtain a comprehensive resource score for multiple cross-language resources, which comprehensively measures the overall suitability of the resources.

[0140] It should be noted that the comprehensive resource score integrates evaluation results from two dimensions: content matching and domain relevance. It can objectively reflect the comprehensive fit between cross-language resources and user needs, providing a scientific basis for resource recommendations.

[0141] For example, when the content matching score and domain relevance score of a cross-language resource are both relatively high values ​​in their respective dimensions, a higher overall resource score can be obtained after combining the set weights.

[0142] In this implementation, multiple cross-language resources can be sorted in descending order of their comprehensive resource scores, and the sorted cross-language resources can be recommended to the user.

[0143] This implementation method acquires resource domain classifications, resource reference relationships, and domain cooperation data for multiple cross-language resources. It then performs semantic alignment between user search requirements and these cross-language resources to determine search intent and semantic alignment tags. A structure is then constructed based on the semantic alignment tags and resource domain classifications to obtain a cross-domain search index for the multiple cross-language resources. Subsequently, the search intent and the cross-domain search index are matched and calculated to obtain a content matching score for the multiple cross-language resources. This score, combined with corresponding weights, is incorporated into the overall resource score, enhancing the accuracy of matching cross-language resources with user search intent, reducing semantic bias in cross-language searches, and making recommended resources more aligned with users' actual needs.

[0144] This approach allows for in-depth exploration of potential domain connections between cross-language resources, breaking through the limitations of evaluating resources based on a single content dimension. It recommends resources with greater domain reference value to users, satisfying their deeper domain exploration needs. By integrating the dual evaluation dimensions of content matching and domain relevance, it balances the relevance of resource content with the value of domain relevance, making the ranking of resource recommendations more scientific and reasonable, and effectively improving users' recognition and satisfaction with the recommended resources.

[0145] Figure 9 A flowchart illustrating the fifth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 9 As shown, in some implementations, in the above-mentioned S210, the semantic alignment labels and resource domain classifications are structured to determine the cross-domain retrieval index for multiple cross-language resources, including S211 to S212. S211 to S212 will be explained in detail below.

[0146] S211. Perform feature extraction on professional documents and structured data to obtain professional element features. Among them, professional element features include professional clause numbers, data fields, and core keywords.

[0147] In this implementation, feature extraction operations can be performed on professional documents and structured data to extract a set of key features with professional identification attributes from the two types of data sources, thereby obtaining professional element features.

[0148] It should be noted that the professional element features cover three core categories: professional clause number, data field, and core keywords. These features can accurately anchor the core information in professional documents and structured data.

[0149] For example, legal clause numbers, case data fields, and core legal keywords can be extracted from legal professional documents and structured case data in the field of law to form corresponding professional element features.

[0150] This implementation method can extract professional element features including professional clause numbers, data fields, and core keywords, which can provide reliable basic data for subsequent knowledge graph construction, making subsequent search indexes more professional and targeted, and effectively improving search accuracy.

[0151] S212. Construct a professional knowledge graph by analyzing the characteristics of professional elements. Then, construct a unified index for the professional knowledge graph to obtain a cross-domain retrieval index.

[0152] In this implementation, a professional knowledge graph can be constructed based on the extracted professional element features.

[0153] It should be noted that this construction process can integrate the characteristics of scattered professional elements according to the inherent logical connections of their respective fields, forming a systematic knowledge network structure with hierarchical relationships and related links.

[0154] For example, the extracted legal professional elements can be used to construct a legal professional knowledge map containing legal nodes, clause nodes, and case data nodes according to legal categories and case application logic.

[0155] In this implementation, a professional knowledge graph is constructed based on the characteristics of professional elements, which can systematically integrate scattered professional information, avoid the fragmentation problem of index construction, and make the domain association of cross-language resources more systematic.

[0156] In this implementation, a unified index building operation can be performed on the completed professional knowledge graph to generate a cross-domain retrieval index.

[0157] It should be noted that the unified index construction can standardize the coding of each node and related link in the professional knowledge graph, providing a unified retrieval identification system for cross-domain retrieval of cross-language resources.

[0158] For example, standardized index codes can be assigned to all nodes in the legal knowledge graph to establish a corresponding mapping relationship between nodes and indexes, forming a cross-domain retrieval index that can support cross-language legal resource retrieval.

[0159] This implementation method extracts features from professional documents and structured data to obtain professional element features including professional clause numbers, data fields, and core keywords. Then, a professional knowledge graph is constructed based on these professional element features. Finally, a cross-domain retrieval index is constructed by uniformly indexing the professional knowledge graph, making the construction of the cross-domain retrieval index more professional and structured. It can more accurately associate cross-language resources with domain classification and semantic alignment tags, effectively improving the accuracy of retrieval.

[0160] This implementation avoids the fragmentation problem in index building, making the domain classification and semantic alignment tag association of cross-language resources more systematic, improving the accuracy of subsequent content matching scores and domain association scores; it also improves the processing efficiency of retrieval and recommendation, and makes the recommended resources more in line with the user's in-depth professional needs, effectively reducing retrieval bias.

[0161] Figure 10 A flowchart illustrating the sixth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 10 As shown, in some implementations, the above method also includes S310 to S320, which will be described in detail below.

[0162] S310. Obtain historical search risk cases where resource search results and user search needs do not match. Extract risk dimensions from these historical search risk cases to determine their domain risk dimension characteristics. Define rules for these domain risk dimension characteristics and determine the risk dimension weight rules. Historical search risk cases include professional document search results, structured data search results, and user search needs.

[0163] Figure 11 A schematic diagram illustrating the workflow of the sixth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 11 As shown, this implementation can systematically collect case information where resource search results and user search needs do not match in past search scenarios, providing a solid sample foundation for subsequent risk dimension analysis.

[0164] It should be noted that historical search risk cases consist of professional document search results, structured data search results, and user search needs. The combination of these three types of data can fully present the actual scenarios of search mismatch.

[0165] For example, cases in the biomedical field where professional literature search results deviate from the user's target research needs can be collected and included in the historical search risk case database.

[0166] In this implementation, multi-dimensional feature extraction can be carried out on the collected historical search risk cases. From the search logic, resource attributes, and demand description of the cases, domain risk dimension features that can reflect the reasons for search mismatch can be extracted.

[0167] For example, risk dimension features such as "misunderstanding of demand due to ambiguity of professional terms" and "misclassification of resource fields" can be extracted from historical cases to clarify the core causes of search mismatch.

[0168] For example, taking the field of patent document retrieval as an example, four core risk dimensions of the field are extracted from historical retrieval risk cases in this field (such as mismatched cases where the user's search requirement is "control method of thermal management system for new energy vehicle battery", but the search results are mostly professional documents on battery material preparation, structured data without core control logic, and patent documents that are more than 10 years old). Then, rules are defined for each dimension.

[0169] For example, the topic relevance dimension rule is as follows: if the overlap between the core technical topic and the user's demand topic in the search results is less than 60%, it is considered high risk; 60%-80% is considered medium risk; and above 80% is considered low risk. The field matching accuracy dimension rule is as follows: if the number of matches between core fields such as the patent's IPC classification number and claim keywords and user demand keywords in the structured data is less than 50% of the total number of demand fields, it is considered high risk; 50%-70%... The risk level is determined as medium, with over 70% classified as low. For the timeliness of results, results with patent publication dates more than 10 years from the search date are considered high-risk, 5-10 years as medium-risk, and less than 5 years as low-risk. For the authority of the source, results not from official patent databases such as the State Intellectual Property Office are considered high-risk, those from official databases without corresponding patent families are considered medium-risk, and those from official databases with multiple patent families from various countries are considered low-risk. Furthermore, weighting rules are determined based on the priority of each dimension's impact on the matching degree between search results and user needs. The theme relevance dimension, which plays a decisive role in core matching, receives the highest weight of 40%; the field matching accuracy dimension, which supports accurate matching of structured data, receives 25%; the result timeliness dimension, which aligns with the technological iteration characteristics of the new energy vehicle field, receives 20%; and the source authority dimension, which ensures the credibility of search results, receives 15%. The total weight of each dimension is 100%.

[0170] S320. Determine the risk dimension scores for resources across multiple domains using risk dimension weighting rules. Obtain the matching recommendation weights and risk dimension weights. Determine the difference between the product of the matching recommendation value and the matching recommendation weight for each resource across multiple domains, and the product of the risk dimension score and the risk dimension weight, as the adjusted matching recommendation value for each resource across multiple domains. Recommend resources across multiple domains to the user in descending order of their adjusted matching recommendation values.

[0171] In this implementation, based on the extracted domain risk dimension features and the probability impact of each dimension causing search mismatch, corresponding weight allocation rules can be formulated to clarify the proportion of different risk dimensions in the overall risk assessment.

[0172] For example, a relatively higher weighting is assigned to the "technical term ambiguity" dimension, which has a higher probability of causing search mismatch, thereby strengthening the influence of this dimension in risk assessment.

[0173] In this implementation, the established risk dimension weighting rules can be used to conduct risk assessments on each of the candidate domain resources, determine the score for each risk dimension corresponding to each domain resource, and summarize to obtain the overall risk dimension score.

[0174] For example, resources in a certain field can be substituted into the risk dimension weighting rules, and the scores of the resource in dimensions such as "terminology ambiguity" and "classification bias" can be calculated respectively. The scores are then summed to obtain the risk dimension score.

[0175] In this implementation, preset matching recommendation weights and risk dimension weights can be retrieved according to the needs of the current retrieval scenario. These two types of parameters are used to balance the proportion of matching degree and risk prevention in the recommendation decision.

[0176] For example, in technical retrieval scenarios where high accuracy is required, a relatively higher matching recommendation weight can be configured to ensure the core matching of recommended resources.

[0177] In this implementation, the product of the matching recommendation value and the matching recommendation weight, and the product of the risk dimension score and the risk dimension weight can be calculated according to the established calculation logic. Then, the former is subtracted from the latter to obtain the corrected matching recommendation value for each domain resource.

[0178] For example, the basic matching score is obtained by multiplying the matching recommendation value of a resource in a certain domain by the matching recommendation weight, and the risk dimension score is obtained by multiplying the risk dimension weight by the risk dimension score. The difference between the two is the corrected matching recommendation value of the resource.

[0179] In this implementation, the corrected matching recommendation values ​​of all candidate domain resources can be sorted, and the resource list can be organized in descending order and pushed to the corresponding users.

[0180] This implementation method obtains historical search risk cases where resource search results do not match user search needs. Risk dimensions are extracted from these cases to determine corresponding domain risk dimension features. Rules are then defined for these domain risk dimension features to form risk dimension weight rules. These rules determine risk dimension scores for multiple domain resources. Combining the matching recommendation weight and the risk dimension weight, the difference between the product of the matching recommendation value and the matching recommendation weight and the product of the risk dimension score and the risk dimension weight is calculated to obtain a corrected matching recommendation value. Domain resources are then recommended in descending order of this value. By incorporating search risk factors, domain resources with potential matching biases can be filtered out, improving the relevance of recommended content to user needs and reducing mismatches.

[0181] This implementation method leverages risk experience from historical searches to build targeted risk control logic, enabling resource recommendations to consider not only matching accuracy but also past search biases, thus improving the reliability and rationality of resource recommendations. It also enriches the decision-making dimensions of resource recommendations, shifting from a single matching accuracy consideration to a dual consideration of matching accuracy and risk control, enabling the recommendation of domain resources that better meet users' actual needs and avoid search risks, thereby optimizing the overall search and recommendation experience.

[0182] Figure 12 A flowchart illustrating the seventh intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 12 As shown, in some implementations, the above method also includes S330 to S340, which will be described in detail below.

[0183] S330. Determine the sum of the products of structural matching score and structural matching weight, and semantic matching score and semantic matching weight for multiple cross-language resources, as the comprehensive resource score for multiple cross-language resources.

[0184] In this implementation, the structural matching scores and corresponding structural matching weights of multiple cross-language resources can be multiplied together, the semantic matching scores and corresponding semantic matching weights can be multiplied together, and the two sets of product results can be added together to obtain the comprehensive resource score of multiple cross-language resources.

[0185] It should be noted that the comprehensive resource score takes into account the degree to which cross-language resources match users' search needs at both the structural and semantic levels. By balancing the influence of the two dimensions through weighting, it can objectively reflect the matching value of resources.

[0186] S340. Determine the difference between the product of the overall resource score, risk dimension score, and risk dimension weight of multiple cross-language resources, and use this difference as the revised overall resource score for the multiple cross-language resources. Recommend multiple cross-language resources to the user in descending order of the revised overall resource score.

[0187] In this implementation, the risk dimension scores and corresponding risk dimension weights of multiple cross-language resources can be calculated first, and then the product of the obtained comprehensive resource score can be subtracted to obtain the corrected comprehensive resource score of multiple cross-language resources.

[0188] It should be noted that this difference calculation method introduces a negative correction for risk factors, which can appropriately reduce the score of resources with potential search risks and filter out resources with high matching degree but with potential risks.

[0189] For example, if a cross-language resource has a high overall resource score, but the corresponding risk dimension score shows that it has had a search risk of deviation in the expression of professional terms, this correction method can be used to reduce its revised overall resource score.

[0190] In this implementation, multiple cross-language resources can be sorted in descending order of their comprehensive modified resource scores, and then the sorted cross-language resources can be pushed to the user's end to present a resource list that meets the user's needs.

[0191] This implementation first determines the sum of the product of the structural matching score and the structural matching weight, and the product of the semantic matching score and the semantic matching weight of the cross-language resources, as the comprehensive resource score of the cross-language resources. Then, the product of the risk dimension score and the risk dimension weight is subtracted from this comprehensive resource score to obtain the corrected comprehensive resource score. Finally, cross-language resources are recommended to users in descending order of the corrected comprehensive resource scores. This approach retains the matching degree between cross-language resources and user search needs while incorporating negative corrections for risk factors, making the recommendation results more in line with the user's actual needs while avoiding potential search risks.

[0192] This implementation addresses the shortcomings of previous methods that failed to consider search risks, filters out cross-language resources with high matching degrees but retrieval risks, and improves the reliability and accuracy of recommendation results. It not only ensures the matching degree between cross-language resources and user search needs, but also reduces invalid or risky recommended content through risk correction, optimizes the user search experience, and improves user satisfaction with search results.

[0193] Figure 13 A flowchart illustrating the eighth intelligent text retrieval and analysis method integrating machine language and natural language provided in this application embodiment is shown below. Figure 13 As shown, in some implementations, the above method also includes S350 to S360, which will be described in detail below.

[0194] S350. Determine the sum of the product of the content matching score and content matching weight of multiple cross-language resources, and the product of the domain association score and domain association weight, as the comprehensive resource score of multiple cross-language resources.

[0195] In this implementation, a multiplication operation can be performed on the content matching scores and corresponding content matching weights of multiple cross-language resources, and a multiplication operation can be performed on the domain association scores and corresponding domain association weights. The results of the two operations are then summed to obtain the comprehensive resource score of multiple cross-language resources.

[0196] It should be noted that this calculation method can simultaneously consider the content relevance of cross-language resources and user search intent, as well as the relevance value of resources within the domain. By balancing the influence of the two dimensions through weight allocation, the overall resource score can better reflect actual search needs.

[0197] S360. Determine the difference between the product of the overall resource score, risk dimension score, and risk dimension weight of multiple cross-language resources, and use this difference as the revised overall resource score for the multiple cross-language resources. Recommend multiple cross-language resources to the user in descending order of the revised overall resource score.

[0198] In this implementation, the comprehensive resource score of multiple cross-language resources can be subtracted from the product of the risk dimension score and the corresponding risk dimension weight, and the difference obtained can be used as the corrected comprehensive resource score of multiple cross-language resources.

[0199] It should be noted that this correction mechanism introduces risk dimension indicators extracted from historical search risk cases to negatively adjust the overall resource score. This can effectively filter out cross-language resources with high matching degree but with search bias risk, thereby improving the reliability of recommendation results.

[0200] For example, for cross-language resources that have previously presented risks related to ambiguous searches using technical terms, the overall score of the corrected resource can be reasonably reduced by multiplying and deducting the risk dimension weights.

[0201] In this implementation, cross-language resources can be pushed to users in descending order of their comprehensive scores based on the modified resource ratings.

[0202] This implementation first determines the content matching score and domain relevance score of cross-language resources. Then, the content matching score is multiplied by its weight, and the domain relevance score is multiplied by its weight. The sum of these two products is used as the overall resource score for the cross-language resource. Next, the product of the risk dimension score and its weight is subtracted from this overall resource score to obtain a revised overall resource score. Cross-language resources are then recommended to users in descending order of their revised overall resource scores. This approach more comprehensively considers the matching degree between cross-language resources and the user's search intent, as well as the domain relevance. Combined with risk dimension correction, this effectively improves the accuracy and reliability of the recommended resources.

[0203] This implementation effectively avoids the risk of resource and demand mismatch in historical searches, reduces errors in recommendation results, and makes recommended cross-language resources more in line with users' actual needs. It deeply mines users' search intent, takes into account the domain relevance of resources, and filters out unsuitable resources through risk correction, significantly improving the adaptability and practicality of cross-language resource recommendations.

[0204] In some implementations, in S120 above, multiple attractiveness evaluation indicators for resources in multiple fields are determined based on search demand tags, historical search data, and comprehensive demand data for resources in multiple fields, including S121 to S122. S121 to S122 will be explained in detail below.

[0205] S121. Determine the text similarity between the search requirement tags and the tags of resources in multiple domains, and determine the historical search similarity between historical search data and historical search data of resources in multiple domains. Obtain the resource timeliness score, resource authority score, resource completeness score, and user preference score for resources in multiple domains. Among them, the user preference score is used to characterize the score corresponding to the click-through conversion rate and average dwell time percentage of domain resources.

[0206] In this implementation, the semantic features of the search requirement tags and the tag texts of resources in various fields can be compared to quantify the degree of matching between the two and obtain the tag text similarity.

[0207] In this implementation, resource selection features and user behavior features from historical search data can be extracted and matched with the corresponding features of current resources in various fields to obtain the historical search similarity.

[0208] In this implementation, the resource timeliness score of each domain resource can be calculated from the time dimension by extracting information such as the release time and the most recent update time of the domain resources.

[0209] In this implementation, the credibility score of resources in each field can be calculated from the perspective of credibility by verifying the qualifications of the publishing entity of the resources in the field and the number of citations in the industry.

[0210] In this implementation, the resource integrity score of each domain resource can be calculated from the content dimension by verifying information such as the content coverage and the completeness of key information of the domain resources.

[0211] It should be noted that the user preference score is used to characterize the click-through rate and average dwell time ratio of domain resources, and comprehensively reflects the actual acceptance of resources by users.

[0212] In this implementation, user click conversion data and average browsing time per user percentage data for resources in various fields can be collected to calculate the corresponding user preference score.

[0213] For example, for e-commerce product category resources, user preference scores can be mainly based on the product click-through rate and the average time users spend browsing product detail pages; the higher the value, the higher the score.

[0214] S122. The similarity of tag texts, historical search similarity, resource timeliness score, resource authority score, resource completeness score, and user preference score of resources in multiple domains are used as multiple attractiveness evaluation indicators for resources in multiple domains.

[0215] In this implementation, the calculated tag text similarity, historical retrieval similarity, and the obtained resource timeliness score, resource authority score, resource completeness score, and user preference score can be used together as a multi-dimensional evaluation index to measure the attractiveness of resources in various fields.

[0216] In this implementation, by constructing multi-dimensional evaluation indicators, domain resources can be comprehensively evaluated from the perspectives of demand matching, resource quality, and user preferences, thus avoiding the limitations of single-dimensional evaluation.

[0217] In some implementations, S120 above involves obtaining the weights of multiple evaluation indicators for resources in each domain. The product of the attractiveness evaluation indicators and the weights of these indicators for the resources in multiple domains is then determined as the matching recommendation value for the resources in multiple domains, including S123 and S124. S123 and S124 are explained in detail below.

[0218] S123. Obtain the tag text similarity weight, historical retrieval similarity weight, resource timeliness weight, resource authority weight, resource completeness weight, and user preference weight for resources in various fields.

[0219] In this implementation, corresponding evaluation index weights can be configured for tag text similarity, historical search similarity, resource timeliness score, resource authority score, resource completeness score, and user preference score, based on the type of the current search scenario and the user's core needs.

[0220] For example, in the case of professional and technical search scenarios, the weight of resource authority and resource completeness can be increased; in the case of hot information search scenarios, the weight of resource timeliness and user preference can be increased.

[0221] S124. Determine the sum of the following products for resources in multiple domains: the product of tag text similarity and tag text similarity weight, the product of historical retrieval similarity and historical retrieval similarity weight, the product of resource timeliness score and resource timeliness weight, the product of resource authority score and resource authority weight, the product of resource completeness score and resource completeness weight, and the product of user preference score and user preference weight. Use this sum as the matching recommendation value for resources in multiple domains.

[0222] In this implementation, the first weighted value is obtained by multiplying the similarity of the tag texts of resources in each domain by their corresponding weights; the second weighted value is obtained by multiplying the historical search similarity by their corresponding weights; and the third weighted value is obtained by multiplying the resource timeliness score by its corresponding weight. The fourth weighted value is obtained by multiplying the resource authority score by its corresponding weight; the fifth weighted value is obtained by multiplying the resource completeness score by its corresponding weight; and the sixth weighted value is obtained by multiplying the user preference score by its corresponding weight.

[0223] In this implementation, the first to sixth weighted values ​​can be summed to obtain the matching recommendation value of resources in each field. This value can objectively reflect the degree of fit between resources and user needs.

[0224] In this implementation, the matching recommendation value is calculated by weighted summation, which can flexibly adjust the importance of each dimension according to different search scenarios to ensure the relevance and rationality of the recommendation results.

[0225] This implementation uses tag text similarity, historical search similarity, resource timeliness score, resource authority score, resource completeness score, and user preference score as attractiveness evaluation indicators. After obtaining the weights corresponding to each indicator, the matching recommendation value is obtained by calculating the sum of the products of each indicator and its corresponding weight. This approach takes into account the matching degree of search needs, resource quality, and user preferences, thereby significantly improving the accuracy of recommendation results.

[0226] This implementation incorporates user preference scores corresponding to click-through rates and average dwell time percentages of domain-specific resources into attractiveness evaluation metrics. It also combines historical search similarity to mine users' historical behavioral tendencies, along with corresponding weight parameters and matching recommendation values, thereby enhancing the alignment between recommendation results and users' actual usage habits. This effectively improves users' acceptance of recommended resources and their satisfaction with the resources.

[0227] This application also provides an intelligent text retrieval and analysis system that integrates machine language and natural language, including units for implementing the method described above.

[0228] Figure 14 A schematic diagram of the logical structure of an intelligent text retrieval and analysis system integrating machine language and natural language, provided in an embodiment of this application, is shown below. Figure 14 As shown, the system 1 of this embodiment includes a processing unit 11, a storage unit 12, and a transceiver unit 13. The processing unit 11 is used to process data, the storage unit 12 is used to store data, and the transceiver unit 13 is used to send and receive data. The processing unit 11, the storage unit 12, and the transceiver unit 13 cooperate with each other to implement the above-described method. The beneficial effects of the embodiments of this application have been described in the above-described method and will not be repeated here.

[0229] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0230] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0231] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0232] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0233] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0234] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0235] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0236] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An intelligent text retrieval and analysis method integrating machine language and natural language, characterized in that, The method includes: Acquire user search needs and historical search data; determine search need tags based on user search needs; perform searches based on search need tags and historical search data to determine comprehensive demand data for resources in multiple fields; wherein, user search needs are represented in natural language, and search need tags are represented in machine language; Based on search demand tags, historical search data, and comprehensive demand data from multiple domains, we determine multiple attractiveness evaluation indicators for resources in multiple domains; we obtain the weights of multiple evaluation indicators for each domain; we determine the sum of the products of multiple attractiveness evaluation indicators and their weights for multiple domains as the matching recommendation value for multiple domains; and we recommend multiple domain resources to users in descending order of their matching recommendation values. Based on search requirement tags and historical search data, a comprehensive requirement data is determined, including: Based on user search needs and historical search data, search preference tags are determined; disambiguation thresholds are obtained; and the search requirement tags and search preference tags are disambiguated by the rule parsing disambiguation unit configured by the disambiguation thresholds to obtain preliminary search instructions. Among them, search preference tags include high-frequency preference tags and fuzzy query preference tags, and search preference tags and preliminary search instructions are represented by machine language. Obtain the weights of the instruction fields corresponding to the search preference tags; adjust the initial search instructions based on the instruction field weights to determine the structured search instructions; perform a search using the structured search instructions to determine the comprehensive demand data for resources from multiple domains; wherein, the structured search instructions are represented in machine language; Based on search demand tags, historical search data, and comprehensive demand data from multiple domains, several attractiveness evaluation indicators for resources in multiple domains are determined, including: Determine the similarity of search demand tags and tag texts of resources in multiple domains, and determine the historical search similarity of historical search data and resources in multiple domains; obtain resource timeliness score, resource authority score, resource completeness score and user preference score of resources in multiple domains; among them, the user preference score is used to characterize the score corresponding to the click conversion rate and average dwell time of domain resources. The similarity of tag texts, historical search similarity, resource timeliness score, resource authority score, resource completeness score, and user preference score of resources in multiple domains are used as multiple attractiveness evaluation indicators for resources in multiple domains. Obtain the weights of multiple evaluation indicators for resources in each domain; determine the sum of the products of multiple attractiveness evaluation indicators and their weights for resources in multiple domains, as the matching recommendation value for resources in multiple domains, including: Obtain the tag text similarity weight, historical search similarity weight, resource timeliness weight, resource authority weight, resource completeness weight, and user preference weight for resources in various fields; The sum of the following factors is used as the matching recommendation value for resources across multiple domains: the product of tag text similarity and tag text similarity weight, the product of historical retrieval similarity and historical retrieval similarity weight, the product of resource timeliness score and resource timeliness weight, the product of resource authority score and resource authority weight, the product of resource completeness score and resource completeness weight, and the product of user preference score and user preference weight.

2. The method according to claim 1, characterized in that, By using structured search commands, comprehensive demand data for resources across multiple domains can be determined, including: Obtain cross-language resource keywords and resource summaries from multiple cross-language resources; based on the cross-language resource keywords and resource summaries, determine the resource structure features and semantic features of multiple cross-language resources; Based on structured search instructions, structural matching calculations are performed on the resource structure features of cross-language resources to determine the structure matching scores of multiple cross-language resources. Semantic matching calculations are also performed on the semantic features of cross-language resources based on structured search instructions to determine the semantic matching scores of multiple cross-language resources. Structure matching weights and semantic matching weights are then obtained. The sum of the products of the structure matching scores and structure matching weights, and the semantic matching scores and semantic matching weights, is determined as the comprehensive resource score for multiple cross-language resources. Finally, multiple cross-language resources are recommended to the user in descending order of their comprehensive resource scores.

3. The method according to claim 2, characterized in that, The method further includes: Acquire resource domain classifications, resource reference relationships, and domain cooperation data for multiple cross-language resources; perform semantic alignment between user search requirements and multiple cross-language resources to determine search intent and semantic alignment tags; construct a structure for semantic alignment tags and resource domain classifications to determine cross-domain search indexes for multiple cross-language resources; The search intent and cross-domain search index are matched and calculated to obtain the content matching score of multiple cross-language resources; the search intent and the resource reference relationship and domain cooperation data of multiple cross-language resources are correlated and quantified to obtain the domain correlation score of multiple cross-language resources. Obtain content matching weight and domain association weight; determine the sum of the product of content matching score and content matching weight, and the product of domain association score and domain association weight of multiple cross-language resources, as the comprehensive resource score of multiple cross-language resources; recommend multiple cross-language resources to users in descending order of comprehensive resource score.

4. The method according to claim 3, characterized in that, The semantic alignment tags and resource domain classifications are structured to determine cross-domain retrieval indexes for multiple cross-language resources, including: Feature extraction is performed on professional documents and structured data to obtain professional element features; among them, professional element features include professional clause numbers, data fields, and core keywords; A professional knowledge graph is obtained by constructing a graph of professional element characteristics; a cross-domain retrieval index is obtained by constructing a unified index of the professional knowledge graph.

5. The method according to claim 4, characterized in that, The method further includes: The process involves: acquiring historical search risk cases where resource search results and user search needs do not match; extracting risk dimensions from these historical search risk cases to determine their domain risk dimension characteristics; defining rules for these domain risk dimension characteristics and determining risk dimension weight rules; and identifying historical search risk cases that include professional document search results, structured data search results, and user search needs. Risk dimension weighting rules are used to determine risk dimension scores for resources in multiple domains; matching recommendation weights and risk dimension weights are obtained; the difference between the product of the matching recommendation value and the matching recommendation weight for resources in multiple domains, and the product of the risk dimension score and the risk dimension weight, is used as the corrected matching recommendation value for resources in multiple domains; resources in multiple domains are recommended to users in descending order of their corrected matching recommendation values.

6. The method according to claim 5, characterized in that, The method further includes: The sum of the products of structural matching score and structural matching weight, and semantic matching score and semantic matching weight, for multiple cross-language resources is used as the comprehensive resource score for multiple cross-language resources. The difference between the product of the overall resource score, risk dimension score, and risk dimension weight of multiple cross-language resources is determined as the corrected overall resource score for the multiple cross-language resources; the multiple cross-language resources are recommended to the user in descending order of the corrected overall resource score.

7. The method according to claim 6, characterized in that, The method further includes: The sum of the product of content matching score and content matching weight, and the product of domain association score and domain association weight of multiple cross-language resources is determined as the comprehensive resource score of multiple cross-language resources; The difference between the product of the overall resource score, risk dimension score, and risk dimension weight of multiple cross-language resources is determined as the corrected overall resource score for the multiple cross-language resources; the multiple cross-language resources are recommended to the user in descending order of the corrected overall resource score.

8. An intelligent text retrieval and analysis system integrating machine language and natural language, characterized in that, Includes units for implementing the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Low-resource language open domain question answering method based on multi-dimensional answer screening

    CN120450047A

  • Time dimension semantic retrieval method based on large model

    CN121301383A