Data networking search sorting method and system based on data pragmatics
By combining the sorting method of query similarity and pragmatic weight values in the digital network search system, the problem of insufficient search accuracy is solved, and more comprehensive search coverage and user needs matching are achieved, and the system's accuracy rate and user experience are improved.
Patent Information
- Application Number
- CN202510838708.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing digital network search system has insufficient retrieval accuracy under the security requirements of not leaving the domain, which affects the user experience.
By determining the search result data set that matches the search text in the registry of the digital network, and sorting it in combination with the query similarity and pragmatic weight values of the digital objects, the search results range is expanded and the pragmatic relationship between digital objects is considered.
It improves the accuracy of search results and the efficiency of information acquisition, ensures that the sorting results meet user needs, and improves the accuracy rate and user satisfaction of the Digital Network search system.
Smart Images

Figure CN120354018A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of digital network search, and particularly relates to a digital network search ranking method, system, device, and storage medium based on data pragmatics. Background Art
[0002] Data of various industries and organizations is stored locally to form data spaces. Data within a data space can be shared and circulated, and data between data spaces has security requirements that it cannot leave the domain. Based on the digital network, data can be encapsulated into digital objects (Digital Object, DO), supporting data retrieval and access through digital object identifiers and metadata. The query of digital objects depends on the interaction between the digital object registry and the digital object repository system (Digital Object Repository, DORP), and data users complete search requests through the digital object interface protocol (Digital Object Interface Protocol, DOIP). The security requirement that data in the data space does not leave the domain makes the traditional search engine search framework based on the processes of data aggregation, indexing, and retrieval no longer applicable.
[0003] The current digital network search system mainly filters search results based on inverted index and query similarity calculation.
[0004] However, under the security requirement of not leaving the domain, there is a problem of insufficient retrieval accuracy in the process of retrieving data objects, which affects the user experience. Summary of the Invention
[0005] This application aims to provide a digital network search ranking method, system, device, and storage medium based on data pragmatics, which solves the problems of insufficient retrieval accuracy of data objects and poor retrieval experience on the basis of meeting the requirements that data in each data space does not leave the domain.
[0006] In a first aspect, an embodiment of this application discloses a digital network search ranking method based on data pragmatics, including: Determine a retrieval result data set that matches the retrieval text used to describe the target digital object to be retrieved from the registries of the digital network distributed in multiple dispersed data spaces; the retrieval result data set includes a first digital object that matches the retrieval text and a second digital object that is pragmatically associated with the first digital object; Determine the query similarity of the metadata of each digital object in the retrieval result data set to the retrieval text, and determine the pragmatic weight value of each digital object in the retrieval result data set; Determine the retrieval result evaluation value of each digital object in the retrieval result dataset according to the query similarity and pragmatic weight value of each digital object in the retrieval result dataset, so as to obtain the sorting result of the digital objects in the retrieval result dataset.
[0007] In a second aspect, an embodiment of the present application further discloses a digital network search sorting system based on data pragmatics, including: A dataset module, configured to determine a retrieval result dataset that matches a retrieval text for describing a target digital object to be retrieved from a registry of a digital network distributed in multiple scattered data spaces; the retrieval result dataset includes a first digital object that matches the retrieval text and a second digital object that is pragmatically associated with the first digital object; A parameter calculation module, configured to determine the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text respectively, and determine the pragmatic weight value of each digital object in the retrieval result dataset; An evaluation and sorting module, configured to determine the retrieval result evaluation value of each digital object in the retrieval result dataset according to the query similarity and pragmatic weight value of each digital object in the retrieval result dataset, so as to obtain the sorting result of the digital objects in the retrieval result dataset.
[0008] In a third aspect, an embodiment of the present application further discloses an electronic device, including a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] In summary, in the embodiments of the present application, the retrieval result data set composed of digital objects obtained from different data spaces not only includes the first digital objects directly matching the retrieval text, but also extends to the second digital objects pragmatically associated with the first digital objects, enabling the search results to no longer be limited to text matching. Instead, based on the pragmatic relationships between the digital objects stored in different data spaces, without the data in each data space leaving the domain, a more comprehensive search coverage range is achieved, improving the scenario-based matching degree of the data. Furthermore, by calculating the query similarity of each digital object to evaluate its direct relevance to the retrieval text, it is ensured that the retrieval results can preferentially return objects with a high matching degree to the query text, avoiding the situation where the retrieval results deviate from the true needs of the user. At the same time, a pragmatic weight evaluation mechanism is further introduced, and the weight is calculated according to the pragmatic relevance between digital objects, breaking through the limitation of existing search ranking algorithms that only rely on text matching or web link analysis, enabling the retrieval results to reflect the usage value of digital objects in the data space. This not only optimizes the logic of search ranking, but also makes the ranking results more in line with the needs of users, improving the retrieval accuracy and information acquisition efficiency. Finally, the ranking of the retrieval results is comprehensively calculated using the query similarity and pragmatic weight values, enabling the final ranking results to reflect the relevance of data objects and the importance of application scenarios, enhancing the precision rate and user satisfaction of the digital network search system. Thus, based on the method of the embodiments of the present application, it is ensured that the data ranking matches the actual application relationship, thereby improving the search experience of users in a distributed digital network storage environment. By optimizing the retrieval accuracy, not only is the precision rate of the search system improved, but also the retrieval error caused by unreasonable ranking is effectively reduced, enhancing the convenience for users to obtain key information. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the drawings: Figure 1 is a flowchart of the steps of a digital network search ranking method based on data pragmatics provided by an embodiment of the present application; Figure 2 is a flowchart of the steps of another digital network search ranking method based on data pragmatics provided by an embodiment of the present application; Figure 3 is the generation process of a retrieval result data set under an embodiment of the present application; Figure 4 is the calculation process of a pragmatic weight value under an embodiment of the present application; Figure 5 is a complete ranking process under an embodiment of the present application; Figure 6 is a block diagram of a digital network search ranking device based on data pragmatics provided by an embodiment of the present application; Figure 7It is a block diagram of an electronic device according to an embodiment provided by an embodiment of the present application. Detailed implementation manners
[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0013] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object may be one or multiple. In addition, "and / or" in the present application means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0014] In the present application, considering that data is not only an isolated information storage unit, but also a digital object that is interrelated and has a pragmatic relationship (Data Pragmatics), therefore, as Figure 1 shown, it is a method for number network search and sorting based on data pragmatics provided by an embodiment of the present application.
[0015] The method may include the following steps: Step 101, determine a retrieval result data set that matches the retrieval text used to describe the target digital object to be retrieved from the registry of the number network distributed in multiple dispersed data spaces respectively.
[0016] Among them, the retrieval result data set includes a first digital object that matches the retrieval text, and a second digital object that is pragmatically associated with the first digital object.
[0017] In some embodiments of the present application, since it is difficult to directly match digital objects stored in a dispersed data space relying solely on the retrieved text to meet the query requirements in the digital networking environment, the search system needs to expand the scope of retrieved results to enhance data relevance and search coverage. Therefore, the search system will determine a dataset of retrieved results that match the text based on the retrieved text input by the user. Specifically, the system can first parse the retrieved text input by the user and query in the digital object registry to obtain the first digital object that matches. Subsequently, based on the pragmatic relationship of the first digital object, the system further associates and expands to the second digital object that has a pragmatic association with it to form a complete dataset of retrieved results. The pragmatic relationship (Data Pragmatics) refers to the usage relationship of digital objects in different application scenarios, which determines the association strength and applicability between data objects. After performing this step, the search results are no longer limited to text matching, but can generate a more comprehensive dataset of search results based on the application relationship between data, which reduces the search deviation caused by the current digital networking search system relying solely on text matching and improves the scenario matching degree of the search results.
[0018] In a specific example, the user enters the keyword "distributed storage structure" in the search system, aiming to obtain data resources related to this technology. During the execution process, the search system retrieves the first digital object that matches this keyword in each digital object registry, such as a technical document about the "design specification of distributed storage architecture". Then, the system analyzes the metadata of the first digital object, identifies its pragmatic relationship with other data objects, and expands the dataset of search results to include the associated second digital object, such as a research literature on "distributed storage optimization algorithm", in the retrieved results. In this way, the retrieved results finally obtained by the user not only include the technical document that directly matches the search text, but also cover a series of relevant materials expanded based on the pragmatic relationship, enabling the search results to more accurately meet the research needs of the user and improving the retrieval efficiency and comprehensiveness of information acquisition.
[0019] Step 102: Determine the query similarity of the metadata of each digital object in the retrieved result dataset to the retrieved text respectively, and determine the pragmatic weight value of each digital object in the retrieved result dataset.
[0020] In some embodiments of the present application, considering that solely relying on text matching cannot fully reflect the actual relevance of data objects, it is necessary to combine data pragmatic relationships to determine their pragmatic weight values while calculating the query similarity of each digital object. Therefore, the search system needs to evaluate the digital objects in the retrieved result dataset to ensure that the sorting result meets the requirements of the user's query. During the execution process, the system first parses the retrieved text and extracts query keywords for subsequent similarity calculation. Then, for each digital object in the retrieved result dataset, the system uses the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm to calculate the query similarity of the object to measure the matching degree between its text content and the query text. Subsequently, the system constructs a link structure between digital objects based on data pragmatic relationships and uses the Digital Object Rank (DO-Rank) algorithm to calculate the pragmatic weight value to measure the importance of digital objects in the data space. The pragmatic weight value reflects the influence degree of digital objects in a specific application scenario. In this way, the search system can sort digital objects based on the comprehensive evaluation results, improving the accuracy of the retrieved results.
[0021] In a specific example, the user enters the keyword "distributed storage architecture" in the search system, expecting to obtain relevant technical materials. During the execution process, the system first extracts the keyword "distributed storage" for query similarity calculation and calculates the similarity values of multiple technical documents matching this keyword through the TF-IDF algorithm. Subsequently, the system further analyzes the metadata of each document in the retrieved result dataset, identifies its pragmatic relationships with other digital objects, and uses the DO-Rank algorithm to calculate the pragmatic weight value of each document based on the pragmatic relationship structure. For example, a document describing "distributed storage optimization strategies" may obtain a higher weight due to strong pragmatic relationships, while a document only involving "evaluation of storage device hardware performance" may obtain a lower weight. After completing the execution process, the search results finally obtained by the user not only consider the text matching degree but also incorporate the influence of pragmatic relationships, making the search results more in line with the user's actual needs and improving the retrieval accuracy and information value.
[0022] Step 103: Determine the retrieval result evaluation value of each digital object in the retrieved result dataset according to the respective query similarity and pragmatic weight value of each digital object in the retrieved result dataset, so as to obtain the sorting result of the digital objects in the retrieved result dataset.
[0023] In some embodiments of the present application, considering that solely relying on the matching degree of the query text cannot fully reflect the actual value of data objects, while combining pragmatic relationship evaluation can improve the rationality of sorting and make the sorting result more in line with the actual application scenario of the data, the search system needs to determine the retrieval result evaluation value based on the query similarity and pragmatic weight value of each digital object in the retrieval result dataset, so as to finally generate a sorting result. The system can first calculate the query similarity of each digital object in the retrieval result dataset, and this value is used to measure the matching degree between the text content of the digital object and the query text. Subsequently, the system combines the weighted calculation of the data pragmatic relationship to determine the pragmatic weight value of each digital object, and this value reflects the relevance and application value of the digital object in the data network environment. After the calculation of the query similarity and the pragmatic weight value is completed, the system uses linear weighting or other optimization methods to combine the two to generate the retrieval result evaluation value, and sorts all digital objects based on this evaluation value. In this way, the search system can adjust the sorting of the search results according to the pragmatic attributes of the data, so that the sorting result can not only reflect the text matching degree of the data object, but also comprehensively consider its pragmatic weight, improving the accuracy of retrieval and the effectiveness of data utilization.
[0024] In a specific example, the user enters the keyword "distributed storage architecture" in the search system to obtain relevant technical documents. During the execution process, the system first calculates the query similarity of each document in the retrieval result dataset. For example, a document about "distributed storage optimization algorithm" may have a high similarity, while another document discussing "cloud storage hardware devices" may have a low similarity. Subsequently, the system analyzes the pragmatic relationship between each document and calculates the pragmatic weight value based on the citation frequency, application scenario correlation degree, and weight distribution of the data object. For example, if a certain document is cited in multiple storage system studies, then this document may have a high pragmatic weight. Finally, the system combines the query similarity and the pragmatic weight value, calculates the retrieval result evaluation value, and arranges the digital objects according to the size of the evaluation value. The sorting result finally obtained by the user is not only sorted based on the query text matching degree, but also comprehensively considers the actual application value of the data object, making the search result closer to the user's needs and improving the retrieval efficiency and availability of information.
[0025] In summary, in the embodiments of the present application, the retrieved result data set composed of digital objects obtained from different data spaces not only includes the first digital objects directly matching the retrieved text, but also extends to the second digital objects pragmatically related to the first digital objects, so that the search results are no longer limited to text matching, but can, based on the pragmatic relationships between the digital objects stored in different data spaces, achieve a more comprehensive search coverage without the data in each data space leaving the domain, improving the scenario-based matching degree of the data. Furthermore, by calculating the query similarity of each digital object to evaluate its direct relevance to the retrieved text, it is ensured that the retrieved results can preferentially return the objects with a high matching degree to the query text, avoiding the situation where the retrieved results deviate from the true needs of the user. At the same time, a pragmatic weight evaluation mechanism is further introduced, and the weight is calculated according to the pragmatic relevance between the digital objects, breaking through the limitation of the existing search ranking algorithms that only rely on text matching or web link analysis, enabling the retrieved results to reflect the usage value of the digital objects in the data space, not only optimizing the logic of the search ranking, but also making the ranking results more in line with the needs of the users, improving the retrieval accuracy and information acquisition efficiency. Finally, the ranking of the retrieved results is comprehensively calculated using the query similarity and the pragmatic weight value, so that the final ranking results can reflect the relevance of the data objects and the importance of the application scenarios, enhancing the precision rate and user satisfaction of the digital network search system. Thus, based on the method of the embodiments of the present application, it is ensured that the data ranking matches the actual application relationship, thereby improving the search experience of the user in the distributed digital network storage environment. By optimizing the retrieval accuracy, not only the precision rate of the search system is improved, but also the retrieval error caused by unreasonable ranking is effectively reduced, enhancing the convenience for the user to obtain key information.
[0026] Figure 2 This is another digital network search ranking method based on data pragmatics provided by the embodiments of the present application.
[0027] The method may include the following steps: Step 201, determine a retrieved result data set that matches the retrieved text used to describe the target digital object to be retrieved from the registry of the digital network respectively distributed in multiple scattered data spaces.
[0028] Among them, the retrieved result data set includes the first digital objects matching the retrieved text and the second digital objects pragmatically related to the first digital objects.
[0029] The method shown in this step has been described in step 101 and will not be elaborated here.
[0030] Optionally, step 201 includes the following sub-steps: Sub-step 2011, extract the matching text for the target digital object from the retrieved text.
[0031] In some embodiments of the present application, the search system needs to extract text that matches the target digital object from the retrieved text to ensure that the subsequent retrieval process can correctly locate the data object. The reason for performing this step is that the user's retrieved text usually contains multiple words or phrases, some of which may be directly related to the target digital object, while others may be irrelevant information or modifiers. Therefore, text parsing and screening are required to improve the accuracy of the retrieval. During the execution process, the system first preprocesses the input retrieved text, including word segmentation, stop word removal, punctuation filtering, and stemming, to ensure that the extracted text has a standardized format. Subsequently, the system analyzes the matching degree between the retrieved text and the metadata of the digital objects in the existing data space based on the inverted index or semantic vector model, and filters out the matching text that best conforms to the characteristics of the target digital object. In this way, the search system can improve the retrieval accuracy based on the refined matching text, reduce the interference of invalid data, and enhance the relevance of the search results.
[0032] In a specific example, the user enters "efficient distributed data storage architecture" in the search system to obtain relevant technical materials. During the execution process, the search system first performs word segmentation on the retrieved text to obtain keywords such as "efficient", "distributed", "data storage", and "architecture". Subsequently, the system filters out the modifier "efficient" and analyzes the matching degree between the remaining keywords and the metadata of the digital objects based on the inverted index or semantic vector model. Finally, the system extracts "distributed data storage architecture" as the matching text for subsequent retrieval. After completing the execution process in this way, the user's query can accurately locate the technical documents related to the distributed storage architecture, without being affected by irrelevant modifier words, making the search results more in line with the user's needs and improving the efficiency of information query.
[0033] Sub-step 2012: Match the matching text with the metadata of the digital objects stored in each data space recorded in the registry of each data space, so as to determine the first digital object from at least one data space according to the matched target metadata, and determine the second digital object according to the pragmatic association relationship between the digital objects stored in the data space recorded in the registry corresponding to the determined first digital object.
[0034] In some embodiments of the present application, in order to expand the scope of search results and enhance data relevance and search coverage, the matching text is matched with the metadata of the digital objects stored in each data space recorded in the registry of each data space, so as to determine the first digital object from at least one data space as the basic digital object (BDO) according to the matched target metadata, and then according to the pragmatic association relationship between the digital objects stored in the data space recorded in the registry corresponding to the determined first digital object, determine the second digital object as the extended digital object (EDO) corresponding to the BDO. Specifically, the system can first parse the user's search text and query in the digital object registry to obtain the matching first digital object. Subsequently, based on the pragmatic relationship of the first digital object, the system further associates and extends to the second digital object that has a pragmatic association with it to form a complete search result data set. After performing this step, the search results are no longer limited to text matching, but can generate a more comprehensive search result data set according to the application relationship between data, which reduces the search deviation caused by traditional search systems relying only on text matching and improves the scenario matching degree of search results.
[0035] In a specific example, the user enters the keyword "distributed storage structure" in the search system, aiming to obtain data resources related to this technology. During the execution process, the search system retrieves the first digital object that matches the keyword in the digital object registry, such as a technical document about the "design specification of distributed storage architecture". Then, the system analyzes the metadata of the first digital object, identifies its pragmatic relationship with other data objects, and expands the search result data set to include the associated second digital object, such as a research literature on "distributed storage optimization algorithm", in the search results. In this way, the search results finally obtained by the user not only include the technical documents directly matching the search text, but also cover a series of relevant materials extended based on the pragmatic relationship, enabling the search results to more accurately meet the user's research needs and improving the retrieval efficiency and comprehensiveness of information acquisition.
[0036] Optionally, sub-step 2012 includes the following sub-steps: Sub-step 20121, obtain the registry deployed in each data space, and match the matching text with the metadata of the digital objects in the data space recorded in each registry, so as to determine the first digital object from the data space according to the matched target metadata.
[0037] In some embodiments of the present application, considering that the retrieved text is usually the query content input by the user, while the digital objects in the data space store structured information, the matching relationship between the two needs to be determined through calculation and analysis to ensure the accuracy of the retrieval results. The search system needs to match the matching text with the metadata stored in the digital object registry in the data space to determine the first digital object most relevant to the retrieved text. Specifically, the system first loads the digital object metadata in the data space and performs keyword extraction and text normalization processing on the matching text. Subsequently, the system uses an inverted index or a text similarity calculation method based on the TF-IDF algorithm to match the matching text with the digital object metadata to filter out the target metadata that best meets the query content. In this way, the search system can accurately identify the first digital object in the data space based on the matching text, making the search results highly consistent with the user's needs and improving the accuracy of the retrieval and the query matching quality.
[0038] In a specific example, the user enters the keyword "distributed storage architecture" in the search system to obtain relevant technical documents. During the execution process, the system first loads all the stored metadata from the digital object registry in the data space and performs word segmentation on the input retrieved text to extract core terms such as "distributed storage" and "architecture". Subsequently, the system queries the data space based on the inverted index structure and uses the TF-IDF algorithm to calculate the similarity value between each digital object metadata and the matching text. For example, a document about "distributed storage optimization strategy" may obtain a higher matching degree, while a document discussing "storage device maintenance technology" may have a lower matching degree. Finally, the system filters out the digital object with the highest similarity as the first digital object and uses it for subsequent search sorting. After completing the execution process, the retrieval results obtained by the user can accurately reflect their query needs, enabling the search system to provide more accurate and user-demand-oriented data resources and improving the retrieval quality and search experience.
[0039] Sub-step 20122, determine the second digital object corresponding to the determined first digital object from each registry according to the preset pragmatic depth association relationship.
[0040] As Figure 3 shown, in some embodiments of the present application, considering that the digital objects matched only based on the retrieved text cannot fully reflect the application relationship between the data in the digital network, expanding the pragmatic association objects can optimize the search coverage and improve the scene matching degree of the retrieval. Therefore, the search system needs to identify and expand the EDO ( Figure 3 the arrow in Figure 3 ) that has a pragmatic depth association with the first digital object, that is, the BDO ( Figure 3in E) to determine the second digital object. Specifically, the system first analyzes the metadata of the first digital object to identify its pragmatic relationship with other digital objects. Subsequently, based on the pragmatic weight evaluation mechanism, the system filters out the digital objects that meet the preset pragmatic depth association relationship and identifies them as the second digital objects. The pragmatic depth association relationship (Pragmatic Depth Association) refers to the pragmatic hierarchical relationship of digital objects in a specific application scenario, which determines the pragmatic connectivity and data transferability of search results. For example, Figure 3 the second digital objects in Figure 3 include second digital objects with pragmatic depth association relationships of 1 (second digital objects directly associated with the first digital object) or 2 (second digital objects indirectly associated with the first digital object through other second digital objects) with the first digital object. In this way, the search system can reasonably expand the search scope based on the actual application scenario of the first digital object, so that the sorting results can cover more pragmatically valuable digital objects and improve the application adaptability of the search results.
[0041] In a specific example, the user enters the keyword "distributed storage architecture" in the search system to obtain relevant technical materials. During the execution process, the system first determines that the matching text matches the metadata in the digital object registry and identifies the first digital object, such as a technical document about the "design principles of distributed storage architecture". Subsequently, the system analyzes the pragmatic relationship of the first digital object, identifies its relevance to other digital objects in the data space, and filters out digital objects that meet the preset pragmatic depth, such as a research paper about the "distributed storage optimization algorithm". After completing the execution process, the search results finally obtained by the user not only include directly matching technical documents, but also cover relevant research materials extended based on pragmatic association, making the search results more in line with the user's research needs and improving the retrieval accuracy of the search system and the comprehensive utilization value of information.
[0042] Sub-step 2013, generate a retrieval result data set according to the matched first digital object and second digital object.
[0043] In some embodiments of the present application, the search system needs to generate a retrieval result dataset based on the matched first digital object and second digital object to ensure that the final search results can comprehensively reflect the text matching degree and pragmatic relationship between data objects. The reason for performing this step is that only returning digital objects that directly match the retrieved text may not fully reflect the usage scenarios of the data, and retrieving and expanding by combining pragmatically related objects helps to improve the integrity and usability of the query results. During the execution process, the system first loads the first digital object and its metadata into the retrieval result dataset, and further analyzes the pragmatic association of this object to screen for second digital objects that meet the preset pragmatic depth association relationship. Subsequently, the system merges the first digital object with the second digital object and stores the retrieval result dataset according to the optimized data structure for subsequent sorting and display. The retrieval result dataset refers to the information set containing all relevant digital objects finally formed during the search process, and its structure can support subsequent search optimization and query expansion. After performing this step, the search system can ensure that the search results are not only based on query text matching but also reflect the application associations between data objects, improving the accuracy of retrieval and the comprehensiveness of information acquisition.
[0044] In a specific example, the user enters the keyword "distributed storage architecture" in the search system and obtains relevant technical materials. During the execution process, the system first retrieves the digital object registry to obtain the first digital object that directly matches the retrieved text, such as a technical document about "the design principles of distributed storage architecture". Subsequently, the system analyzes the pragmatic relationship of this first digital object, identifies its associated second digital object, such as a research paper about "distributed storage optimization algorithms", and adds it to the retrieval result dataset. Finally, the system completes the generation of the dataset, enabling the user's retrieval results to include both the technical document that directly matches the query text and the extended materials related to the target topic, improving the richness of the search results and the practical application value of the retrieval.
[0045] Step 202, determine the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieved text respectively, and determine the pragmatic weight value of each digital object in the retrieval result dataset.
[0046] The method shown in this step has been described in step 102 and will not be elaborated here.
[0047] Optionally, in order to determine the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text respectively, step 202 includes the following sub-steps: Sub-step 2021: Extract the retrieval token set of the retrieval text from the retrieval text, and extract the respective metadata token sets from the metadata of the digital objects in each retrieval result dataset.
[0048] In some embodiments of the present application, since the retrieval text usually contains multiple words, some of which may be the core concepts of the query, while other words may be modifying expressions or stop words, it is necessary to parse and filter the text to extract keywords that can effectively represent the query content. Therefore, the search system needs to extract the retrieval token set from the retrieval text to ensure the accuracy of subsequent query similarity calculation. During the execution, the system first performs word segmentation on the input retrieval text, including using natural language processing techniques for word segmentation recognition and removing stop words to ensure that the extracted token set can accurately reflect the query intention. Subsequently, the system performs the same word segmentation on the metadata of the digital objects in each retrieval result dataset to obtain the metadata token set of each digital object, ensuring that the vocabulary expression formats of all data objects are consistent and applicable to subsequent similarity calculations. After executing this step, the search system can unify the text processing method, extract core keywords, improve the accuracy of search matching, and reduce the interference of irrelevant words on the retrieval effect.
[0049] In a specific example, the user inputs the query text "Distributed Storage Architecture Optimization Technology" in the search system to obtain relevant technical documents. During the execution, the system first performs word segmentation analysis on the query text and extracts "Distributed Storage", "Architecture", "Optimization", and "Technology" as the retrieval token set. Subsequently, the system performs the same word segmentation on the metadata of the digital objects in the data space to extract the metadata token set of each document. For example, the metadata token set of a certain technical document may contain keywords such as "Distributed Storage", "Data Management", and "Performance Optimization". After completing this execution process, the search system successfully constructs the token sets of the query text and the metadata, enabling subsequent similarity calculations to be based on a unified vocabulary structure, improving the search accuracy, and ensuring that the retrieval results are highly consistent with the user's query topic.
[0050] Sub-step 2022: Determine the metadata word frequency of each retrieval token in the retrieval token set that appears in each metadata token set respectively, and determine the inverse document frequency of each retrieval token in the retrieval token set that appears in all the metadata token sets respectively.
[0051] In some embodiments of the present application, considering that the metadata of different digital objects may contain the same retrieval terms, but their occurrence frequencies are different, it is necessary to measure the importance of the term in a single digital object and combine the distribution in the entire retrieval result dataset to optimize the text matching degree. The search system needs to calculate the term frequency of each retrieval term in the digital object metadata and the inverse document frequency of the retrieval term in all metadata term sets to ensure the accuracy of the query similarity calculation. During the execution process, the system first scans the metadata term sets of digital objects in each retrieval result dataset, counts the number of occurrences of each term in the retrieval term set in these metadata, and calculates the metadata term frequency (Term Frequency, TF). Subsequently, the system searches for the occurrence of each retrieval term in all metadata term sets and calculates the inverse document frequency (Inverse Document Frequency, IDF). This value reflects the scarcity of the retrieval term in the entire data space. A lower frequency indicates that the term has a higher discrimination degree, while a higher frequency indicates that the term is more common. After performing this step, the search system can adjust the query similarity calculation based on the term frequency and inverse document frequency of the retrieval term, so that terms with higher importance have a greater weight in the final score, improving the accuracy of the search results.
[0052] In a specific example, the user enters the keyword "distributed storage architecture" in the search system to obtain relevant technical documents. During the execution process, the system first traverses the metadata of all digital objects in the retrieval result dataset and counts the occurrence frequencies of "distributed storage" and "architecture" in the retrieval term set in each document. For example, a document about "distributed storage optimization strategy" may contain "distributed storage" multiple times in the metadata, so the metadata term frequency of this term is higher. Subsequently, the system counts the occurrence of the retrieval term in all metadata term sets to calculate the inverse document frequency. For example, if "architecture" appears in multiple documents, its inverse document frequency is lower, while if "distributed storage" appears in only a few documents, its inverse document frequency is higher. After completing this execution process, the search system can reasonably adjust the weights of the term frequency and inverse document frequency, so that the final retrieval similarity calculation can better reflect the discrimination degree of the retrieval term, improving the relevance and matching degree of the search results.
[0053] Sub-step 2023, determine the weighted sum of the metadata term frequency and inverse document frequency of each digital object in each retrieval result dataset for each retrieval term in the retrieval term set as the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text.
[0054] In some embodiments of the present application, since using only word frequency or inverse document frequency alone may cause search results to be biased towards high-frequency words or rare words, and weighted sum calculation can balance the influence of both, enabling the query similarity to accurately reflect the text matching degree. Therefore, the search system needs to calculate the query similarity of each digital object to the retrieved text based on the weighted sum of the metadata word frequency and inverse document frequency of each retrieved token in the retrieved token set, in order to optimize the search matching degree. During the execution process, the system first extracts the token set from the metadata of each digital object in the retrieved result dataset, and calculates the metadata word frequency and inverse document frequency for the retrieved tokens therein. Subsequently, the system performs a weighted sum of the metadata word frequency and inverse document frequency according to a preset weight factor to calculate the query similarity of each digital object. The query similarity is a numerical index that measures the content matching degree between a digital object and the retrieved text, and affects the final sorting result. After executing this step, the search system can calculate and optimize the sorting logic based on the weighted query similarity, making the search results more in line with the user's query requirements and improving the accuracy of retrieval and the effectiveness of information.
[0055] In a specific example, the user enters the keyword "distributed storage architecture" in the search system to obtain relevant technical documents. During the execution process, the system first traverses the metadata of all digital objects in the retrieved result dataset, and calculates the metadata word frequency of "distributed storage" and "architecture" in the retrieved token set. For example, a technical document about "distributed storage optimization strategy" may contain "distributed storage" more frequently, so the metadata word frequency of this word is higher. Subsequently, the system calculates the inverse document frequency of this retrieved token in all metadata token sets. For example, if "architecture" appears in multiple documents, its inverse document frequency is lower, while if "distributed storage" only appears in a few documents, its inverse document frequency is higher. Finally, the system performs a weighted sum of the word frequency and inverse document frequency based on a preset weight factor to calculate the query similarity of each digital object and uses it for subsequent sorting. After completing according to this execution process, the retrieval results finally obtained by the user are not only sorted according to the retrieved text matching degree, but also combined with weighted calculation, making the search results more accurate and improving the accuracy of information query and the search experience.
[0056] Specifically, for the digital object do k to the retrieved text the query similarity can be solved by the formula where represents the word frequency of the token in the query text Q, represents the inverse document frequency.
[0057] Optionally, to determine the pragmatic weight value of each digital object in the retrieval result dataset, step 202 includes the following sub-steps: Sub-step 2024, establish a Markov chain model for characterizing the pragmatic association relationship between digital objects in the retrieval result dataset.
[0058] Among them, the established Markov chain model is used to define the pragmatic weight value between digital objects with a direct association relationship as the association probability value between the digital objects with a direct association relationship in the Markov chain model, and define the weight numerical relationship between the pragmatic weight values between digital objects without a direct association relationship as the probability numerical relationship between the association probability values between digital objects without a direct association relationship in the Markov chain model.
[0059] In some embodiments of the present application, considering that expanding the pragmatic relationship between digital objects in the retrieval result dataset not only affects the sorting of search results, but also determines the application value and scenario matching degree of the data, it is necessary to quantify the association degree between different objects through a probability model. The search system needs to establish a Markov chain model (Markov Chain Model, MCM) to represent the pragmatic association relationship between digital objects in the retrieval result dataset. Specifically, the system first constructs the state space of the Markov chain model and maps all digital objects in the retrieval result dataset to state nodes in the model. Subsequently, the system defines the state transition probability according to the pragmatic weight value of the digital object. Among them, for digital objects with a direct association relationship, their pragmatic weight value is directly defined as the association probability value in the Markov chain, while for digital objects without a direct association relationship, the relationship between their pragmatic weight values is constrained by the numerical relationship between the indirect association probability values. In this way, it can be ensured that the association structure of all data objects can be reasonably reflected in the model. In this way, the search system can accurately depict the pragmatic association relationship of digital objects using the Markov chain model, enabling the sorting logic to not only consider the text matching degree, but also reflect the actual application value of data objects in the digital networking environment, improving the accuracy and rationality of search results.
[0060] In a specific example, a user inputs the keyword "distributed storage architecture" in a search system to obtain relevant technical documents. During the execution, the system first identifies the digital objects in the retrieval result dataset and constructs a Markov chain model, where each digital object is defined as a state node. Subsequently, the system calculates the pragmatic weights between the objects and converts the weight values between directly related digital objects into state transition probabilities. For example, a technical document about "distributed storage optimization strategy" may have a relatively high pragmatic weight and form a high-probability association with the document "performance evaluation of distributed storage architecture". At the same time, the system derives the indirect association probability values based on the weight numerical relationship between non-directly related data objects and establishes appropriate numerical relationships in the model. After completing this execution process, the search system can more accurately represent the pragmatic connections between data objects, making the sorting results more in line with the query requirements and improving the quality of information retrieval and the usability of the search system.
[0061] Sub-step 2025: Solve the Markov chain model to obtain the pragmatic weight value of each digital object.
[0062] In some embodiments of the present application, since the pragmatic relationship between digital objects is not only reflected in the static association structure but also involves dynamic probability transfer, it is necessary to solve the Markov chain to quantify the pragmatic influence of data objects. Therefore, the search system needs to solve the Markov chain model to calculate the pragmatic weight value of each digital object. During the execution, the system first initializes the transition matrix of the Markov chain, which defines the pragmatic association relationship of all digital objects in the retrieval result dataset. Subsequently, the system uses iterative calculation methods, such as the Power Iteration method or the Gauss-Seidel method, to solve the weight distribution of each digital object in the steady state and finally obtain the pragmatic weight value of each digital object. The Markov chain model is used to describe the pragmatic influence of data objects, and reflects the relevance and application value of data through probability transfer relationships. In this way, the search system can optimize the sorting logic based on the solution results of the Markov chain model, enabling data objects with stronger pragmatic influence to obtain higher weights, thereby improving the rationality and pragmatic matching degree of search results.
[0063] In a specific example, a user enters the keyword "distributed storage architecture" in a search system to obtain relevant technical documents. During the execution process, the system first establishes a Markov chain model, where the state nodes correspond to the digital objects in the retrieval result dataset, and initializes the transition probability matrix. For example, a technical document about "distributed storage optimization strategy" may have a relatively high transition probability because it is cited in multiple studies. Subsequently, the system uses the power method for iterative calculation to solve the pragmatic weight values of each data object in the steady state, and finally uses this result for sorting optimization. After completing this execution process, the search results obtained by the user are not only sorted based on the query text matching degree, but also combined with the calculation of pragmatic relationships, making the search results more in line with the user's research needs and improving the retrieval accuracy and the application value of information.
[0064] Optionally, sub-step 2025 includes the following sub-steps: Sub-step 20251: Initialize the pragmatic weight value of each digital object in the retrieval result dataset to a preset initial pragmatic weight value.
[0065] In some embodiments of the present application, considering that the pragmatic weight value is a key parameter affecting search sorting, improper initialization may lead to deviations or instability in the calculation process, thus affecting the rationality of the final retrieval result. Therefore, the search system needs to initialize the pragmatic weight value of each digital object in the retrieval result dataset to ensure that the subsequent iterative calculation can be performed based on a stable initial value. During the execution process, the system first defines a preset initial pragmatic weight value, which can be, for example, 1, and assigns this value to all digital objects in the retrieval result dataset, so that they have a unified pragmatic weight at the beginning of the iterative calculation. The preset initial pragmatic weight value usually adopts a uniform distribution or is set based on prior knowledge to provide a stable input for the Markov chain model. In this way, the search system can ensure that the pragmatic weight calculation of all digital objects starts from a unified initial condition, improve the convergence of the calculation, and provide a reasonable basis for subsequent pragmatic weight optimization.
[0066] In a specific example, a user enters the keyword "distributed storage architecture" in a search system to obtain relevant technical documents. During the execution process, the system first identifies all digital objects in the retrieval result dataset and initializes their pragmatic weight values. For example, an initial weight value of 1 is set for all documents to ensure calculation stability. Subsequently, the system updates the pragmatic weights of each digital object according to the state transition matrix in the Markov chain model to optimize the search sorting. After completing this execution process, the search results obtained by the user are not only sorted based on the query text matching degree, but also combined with the calculation of the initialized pragmatic weights, making the search results more in line with the user's query needs and improving the retrieval accuracy and the rationality of information acquisition.
[0067] In sub-step 20252, update the sum of the products of the ratio of the first pragmatic weight value to the second pragmatic weight value of each third digital object in the retrieval result dataset and the pragmatic weight value of the third digital object, the product of the damping factor coefficient of the retrieval result dataset, and the pragmatic weight evaluation correction value to be the pragmatic weight value of each fourth digital object.
[0068] Among them, the third digital object is the digital object on the upper-level node of the fourth digital object on each Markov chain in the Markov chain model; the first pragmatic weight value is the pragmatic weight value from the third digital object to the fourth digital object on the Markov chain where the third digital object is located; the second pragmatic weight value is the sum of the pragmatic weight values of all digital objects other than from the third digital object to the fourth digital object; the pragmatic weight evaluation correction value is the probability value that the fourth digital object forms a pragmatic association relationship with any digital object in the retrieval result dataset.
[0069] In some embodiments of the present application, due to the dynamic change characteristics of the pragmatic relationship of data objects in the digital networking environment, it is difficult to reflect the actual influence between objects only relying on static weight calculation. Therefore, it is necessary to update the weight value through iteration so that the final sorting can better conform to the application scenario of the data. At this time, the search system needs to update the pragmatic weight value of each fourth digital object to ensure that the Markov chain model can accurately represent the pragmatic association relationship between digital objects. Specifically, considering the following mathematical model: for the digital object (do k ) numbered k, then the calculation method of the pragmatic weight value W k (do prag ) is as shown in the formula: k ), , Among them, DF represents the damping factor, which is used to control the probability of jumping between DOs according to the pragmatic relationship; N represents the total number of all DOs in the current directed graph (that is, the data capacity of the retrieval result dataset); M(do k ) represents the set of digital objects directly pointing to do k ; prag mk represents the pragmatic weight between digital objects do m and do k ; P(do m ) represents the sum of the pragmatic weights of do m pointing to other digital objects. In this way, the search system can optimize the calculation of pragmatic weights based on the Markov chain model, so that the sorting logic not only reflects the text matching degree, but also comprehensively considers the application association between digital objects, improving the rationality and matching degree of search results.
[0070] Such as Figure 4As shown, in the initial state, the initial pragmatic weight values of digital objects do1, do2, do3, and do4 are all 1. The pragmatic weight prag 12 between do1 and do2 is 0.75, the pragmatic weight prag 24 between do2 and do4 is 0.5, the pragmatic weight prag 32 between do2 and do3 is 0.25, and the damping factor DF = 0.85. At this time, if do2 is regarded as the fourth digital object, then the pragmatic weight value W prag of do2 after one round of iteration is calculated according to the following formula: .
[0071] In this way, the search results finally obtained by the user are not only sorted based on the text matching degree, but also comprehensively consider the pragmatic weight optimization, making the search results more in line with the query requirements and improving the accuracy of retrieval and the effectiveness of information acquisition.
[0072] Step 203: Determine the weighted sum of the query similarity and the pragmatic weight value of each digital object in the retrieval result dataset as the retrieval result evaluation value of each digital object in the retrieval result dataset.
[0073] In some embodiments of the present application, since using only the query similarity or the pragmatic weight value alone may lead to sorting deviation, and the method of weighted sum can comprehensively consider the influence of both and improve the rationality of the retrieval result. Therefore, the search system needs to calculate the retrieval result evaluation value of each digital object based on the weighted sum of the query similarity and the pragmatic weight value to ensure that the final sorting can accurately reflect the relevance and application value of the data. Specifically, the system can first extract the query similarity of each digital object from the retrieval result dataset, and this value is used to measure the matching degree between the text content of the digital object and the query text. Subsequently, the system can calculate the pragmatic weight value of the digital object based on the DO-rank algorithm to reflect its pragmatic association strength in the digital network environment. After the calculation is completed, the system uses a predefined weighting factor to perform a weighted sum of the query similarity and the pragmatic weight value, and takes the obtained weighted sum as the retrieval result evaluation value of the digital object. After executing this step, the search system can sort the digital objects in the retrieval result dataset according to this comprehensive evaluation value, so that the final sorting can not only reflect the text matching degree, but also comprehensively consider the pragmatic relationship of the data objects, improving the retrieval accuracy and the information utilization value.
[0074] In a specific example, a user enters the keyword "distributed storage architecture" in a search system to obtain relevant technical documents. During the execution process, the system first calculates the query similarity of each document in the retrieval result dataset. For example, a document about "distributed storage optimization strategies" may have a high query similarity, while another document only related to "general storage hardware analysis" may have a low similarity. Subsequently, the system calculates the pragmatic weight values of these documents. For example, if a certain document is cited in multiple storage system studies, it may obtain a high pragmatic weight. Finally, the system uses a preset weight factor to perform a weighted sum of the query similarity and the pragmatic weight value, such as a query similarity weight of 0.7 and a pragmatic weight of 0.3, to determine the retrieval result evaluation value of each document and complete the sorting of the retrieval results accordingly. After completing according to this execution process, the search results finally obtained by the user are not only sorted based on the query text matching degree but also combined with the pragmatic weight evaluation, making the search results more in line with the actual needs of the user and improving the precision and information value of the retrieval.
[0075] Specifically, the digital object do k The retrieval result evaluation value of the retrieval text Q Specifically, it can be through the formula Record, where Represents the digital object do k The metadata Meta(do k ) and the query similarity between the retrieval text Q, Represents do k 's pragmatic weight value, and α and β are preset weights.
[0076] Step 204, arrange the digital objects according to the magnitude of the retrieval result evaluation value of each digital object in the retrieval result dataset to obtain the sorting result of the digital objects in the retrieval result dataset.
[0077] In some embodiments of the present application, since simply calculating the retrieval result evaluation value cannot directly improve the user's search experience, only through reasonable sorting can it be ensured that high-value data objects are preferentially displayed in the retrieval results and the usability of the retrieval is improved. Therefore, the search system needs to sort the digital objects based on the retrieval result evaluation value of each digital object in the retrieval result dataset to finally determine the arrangement order of the search results. Specifically, the system can first collect the retrieval result evaluation values of all digital objects and arrange them according to a preset sorting rule. For example, descending sorting can be adopted so that the digital objects with higher retrieval result evaluation values are displayed first. The sorting rule can be dynamically adjusted according to the query scenario to optimize the arrangement method of the data objects. In this way, the search system can provide a sorting result comprehensively evaluated based on the query matching degree and pragmatic weight, thereby ensuring that users can preferentially access the most relevant digital objects and improving the efficiency of information query.
[0078] In a specific example, a user enters the keyword "distributed storage architecture" in the search system to obtain relevant technical materials. During the execution process, the system first calculates the retrieval result evaluation values of each document. For example, a document about "distributed storage optimization strategy" may obtain a higher evaluation value, while another document discussing "storage device maintenance technology" may obtain a lower evaluation value. Subsequently, the system arranges these documents in descending order to ensure that the document with the highest evaluation value ranks first in the search results. After completing the execution process according to this, the search results finally seen by the user are sorted according to the comprehensive evaluation of the data objects, so that high-quality data is preferentially displayed, improving the accuracy of the retrieval and the convenience of data acquisition.
[0079] As Figure 5 shown, it is a complete sorting process under the embodiments of the present application: Step S1: The system receives the retrieval text Q input by the user; Step S2: Query based on the digital object registry in multiple data spaces to determine the matching first digital object; Step S3: Based on the pragmatic relationship between digital objects, expand the retrieval scope in multiple data spaces, and determine the second digital object that has a pragmatic association with the first digital object and then query: Step S3.1: Based on the pragmatic relationship between digital objects, expand the retrieval scope to obtain the second digital object, and then obtain the retrieval result dataset; Step S3.2: Calculate the query similarity of each digital object in the retrieval result dataset; Step S3.3: Calculate the pragmatic weight value of each digital object in the retrieval result dataset; Step S3.4: The system performs a final sorting on all digital objects based on the query similarity and the pragmatic weight value, and generates an optimized search result list. This sorting logic makes the data objects with high matching degree and high pragmatic value rank at the front, enabling users to obtain the content that meets their needs more quickly and accurately, and improving the efficiency of information query and the performance of the search system.
[0080] In summary, in the embodiment of the present application, the retrieval result data set composed of digital objects obtained from different data spaces not only includes the first digital objects directly matching the retrieval text, but also extends to the second digital objects pragmatically associated with the first digital objects. This makes the search results not limited to text matching, but can, based on the pragmatic relationships between the digital objects stored in different data spaces, achieve a more comprehensive search coverage without the data in each data space leaving the domain, improving the scenario-based matching degree of the data. Furthermore, by calculating the query similarity of each digital object to evaluate its direct relevance to the retrieval text, it is ensured that the retrieval results can preferentially return the objects with a high matching degree to the query text, avoiding the situation where the retrieval results deviate from the true needs of the user. At the same time, a pragmatic weight evaluation mechanism is further introduced, and the weight is calculated according to the pragmatic relevance between digital objects, breaking through the limitation of the existing search sorting algorithm that only relies on text matching or web link analysis, enabling the retrieval results to reflect the usage value of digital objects in the data space, not only optimizing the search sorting logic, but also making the sorting results more in line with the needs of users, improving the retrieval accuracy and information acquisition efficiency. Finally, the sorting of the retrieval results is comprehensively calculated using the query similarity and the pragmatic weight value, enabling the final sorting results to reflect the relevance of the data objects and the importance of the application scenarios, enhancing the precision rate and user satisfaction of the digital network search system. Thus, based on the method of the embodiment of the present application, it is ensured that the data sorting matches the actual application relationship, thereby improving the search experience of users in a distributed digital network storage environment. By optimizing the retrieval accuracy, not only the precision rate of the search system is improved, but also the retrieval error caused by unreasonable sorting is effectively reduced, enhancing the convenience for users to obtain key information.
[0081] Reference Figure 6 , which shows a digital network search sorting system 30 based on data pragmatics provided by an embodiment of the present application, including: A data set module 301, configured to determine a retrieval result data set that matches the retrieval text for describing a target digital object to be retrieved from the registry of the digital network distributed in multiple scattered data spaces; the retrieval result data set includes a first digital object that matches the retrieval text and a second digital object that is pragmatically associated with the first digital object; The parameter calculation module 302 is used to determine the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text respectively, and determine the pragmatic weight value of each digital object in the retrieval result dataset; The evaluation and sorting module 303 is used to determine the retrieval result evaluation value of each digital object in the retrieval result dataset according to the query similarity and pragmatic weight value of each digital object in the retrieval result dataset, so as to obtain the sorting result of the digital objects in the retrieval result dataset.
[0082] Optionally, the dataset module 301 includes: The retrieval extraction sub-module is used to extract the matching text for the target digital object from the retrieval text; The digital object sub-module is used to match the matching text with the metadata of the digital objects stored in each data space recorded in the registry of each data space, so as to determine the first digital object from at least one data space according to the matched target metadata, and determine the second digital object according to the pragmatic association relationship between the digital objects stored in the data space recorded in the registry corresponding to the determined first digital object; The dataset sub-module is used to generate a retrieval result dataset according to the matched first digital object and second digital object.
[0083] Optionally, the digital object sub-module includes: The first digital object unit is used to obtain the registry deployed in each data space, and match the matching text with the metadata of the digital objects in the data space recorded in each registry, so as to determine the first digital object from the data space according to the matched target metadata; The second digital object unit is used to determine the second digital object corresponding to the determined first digital object from each registry according to the preset pragmatic depth association relationship.
[0084] Optionally, the parameter calculation module 302 includes: The retrieval word segmentation sub-module is used to extract the retrieval word segmentation set of the retrieval text from the retrieval text, and extract the respective metadata word segmentation sets from the metadata of the digital objects in each retrieval result dataset; The metadata word segmentation sub-module is used to determine the metadata word frequency of each retrieval segmentation in each metadata word segmentation set respectively, and determine the inverse document frequency of each retrieval segmentation in all the metadata word segmentation sets respectively; The similarity sub-module is used to determine the weighted sum of the metadata word frequency and the inverse document frequency of each digital object in each retrieval result dataset for each retrieval segmentation in the retrieval word segmentation set as the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text respectively.
[0085] Optionally, the parameter calculation module 302 includes: A modeling sub-module, configured to establish a Markov chain model for characterizing the pragmatic association relationship between digital objects in the retrieval result dataset; the established Markov chain model is used to define the pragmatic weight value between digital objects with a direct association relationship as the association probability value between the digital objects with a direct association relationship in the Markov chain model, and define the weight numerical relationship between the pragmatic weight values between digital objects without a direct association relationship as the probability numerical relationship between the association probability values between digital objects without a direct association relationship in the Markov chain model; A modeling solution sub-module, configured to solve the Markov chain model to obtain the pragmatic weight value of each digital object.
[0086] Optionally, the modeling solution sub-module includes: An initialization unit, configured to initialize the pragmatic weight value of each digital object in the retrieval result dataset to a preset initial pragmatic weight value; An iterative solution unit, configured to update the weighted sum of the ratio of the first pragmatic weight value to the second pragmatic weight value of each third digital object in the retrieval result dataset, the product of the product of the ratio and the pragmatic weight value of the third digital object, the product of the damping factor coefficient of the retrieval result dataset, and the pragmatic weight evaluation correction value as the pragmatic weight value of each fourth digital object; the third digital object is the digital object on the upper-level node of the fourth digital object on each Markov chain in the Markov chain model; the first pragmatic weight value is the pragmatic weight value from the third digital object to the fourth digital object on the Markov chain where the third digital object is located; the second pragmatic weight value is the sum of the pragmatic weight values of all digital objects other than the third digital object to the fourth digital object; the pragmatic weight evaluation correction value is the probability value of the fourth digital object forming a pragmatic association relationship with any digital object in the retrieval result dataset.
[0087] Optionally, the evaluation and sorting module 303 includes: An evaluation value sub-module, configured to determine the retrieval result evaluation value of each digital object in the retrieval result dataset as the weighted sum of the query similarity and the pragmatic weight value of each digital object; A sorting sub-module, configured to arrange the digital objects according to the magnitude of the retrieval result evaluation value of each digital object in the retrieval result dataset to obtain a sorting result of the digital objects in the retrieval result dataset.
[0088] In summary, in the embodiments of the present application, the retrieved result data set composed of digital objects obtained from different data spaces not only includes the first digital objects directly matching the retrieved text, but also extends to the second digital objects pragmatically associated with the first digital objects, so that the search results are no longer limited to text matching, but can be based on the pragmatic relationships between the digital objects stored in different data spaces, and under the premise that the data in each data space does not leave the domain, a more comprehensive search coverage is achieved, and the scenario matching degree of the data is improved; furthermore, by calculating the query similarity of each digital object to evaluate its direct relevance to the retrieved text, it is ensured that the retrieved results can preferentially return the objects with a high matching degree to the query text, avoiding the situation that the retrieved results deviate from the true needs of the user; at the same time, a pragmatic weight evaluation mechanism is further introduced, and the weight is calculated according to the pragmatic relevance between digital objects, breaking through the limitation of the existing search ranking algorithm that only relies on text matching or web link analysis, so that the retrieved results can reflect the usage value of digital objects in the data space, not only optimizing the logic of search ranking, but also making the ranking results more in line with the needs of users, improving the retrieval accuracy and information acquisition efficiency; finally, the ranking of the retrieved results is comprehensively calculated by using the query similarity and the pragmatic weight value, so that the final ranking results can reflect the relevance of data objects and the importance of application scenarios, improving the precision rate and user satisfaction of the digital network search system. Thus, based on the method of the embodiments of the present application, it is ensured that the data ranking matches the actual application relationship, thereby improving the search experience of users in a distributed digital network storage environment. By optimizing the retrieval accuracy, not only the precision rate of the search system is improved, but also the retrieval error caused by unreasonable ranking is effectively reduced, enhancing the convenience for users to obtain key information.
[0089] Referring to Figure 7 , which is a block diagram of an electronic device 500 according to another embodiment of the present invention. For example, the electronic device 500 may be provided as a server. The electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.
[0090] The electronic device 500 includes a processing component 502, which further includes one or more processors, and memory resources represented by the memory 504 for storing instructions executable by the processing component 502, such as application programs. The application programs stored in the memory 504 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 502 is configured to execute instructions to perform the method provided in the embodiments of the present application.
[0091] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components.
[0092] The memory 504 is used to store various types of data to support the operation of the electronic device 500. The memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0093] The power component 506 provides power to various components of the electronic device 500. The power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.
[0094] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera.
[0095] The audio component 510 is used to output and / or input audio signals. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516.
[0096] The input / output I / O interface 512 provides an interface between the processing component 502 and the peripheral interface module.
[0097] The sensor component 514 includes one or more sensors for providing status assessments of various aspects of the electronic device 500. For example, the sensor component 514 can detect the on / off state of the electronic device 500 and the relative positioning of components. The sensor component 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications.
[0098] The communication component 516 is used to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof.
[0099] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.
[0100] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, and the above instructions can be executed by a processor 520 of the electronic device 500 to complete the above methods.
[0101] The electronic device 500 can also operate based on an operating system stored in the memory 504, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM, or the like.
[0102] It should be noted that for the method embodiments of the present application, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequences, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.
[0103] After considering the specification and practicing the application disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
Claims
1. A data pragmatics-based search and sorting method for the data network, characterized in that Including: Determine a retrieval result data set that matches a retrieval text used to describe a target digital object to be retrieved from the registry of the digital network distributed in multiple scattered data spaces respectively; the retrieval result data set includes a first digital object that matches the retrieval text, and a second digital object that is pragmatically associated with the first digital object; Determine the query similarity of the metadata of each digital object in the retrieval result data set to the retrieval text respectively, and determine the pragmatic weight value of each digital object in the retrieval result data set; Determine the retrieval result evaluation value of each digital object in the retrieval result data set according to the query similarity and pragmatic weight value of each digital object in the retrieval result data set respectively, so as to obtain a sorting result of the digital objects in the retrieval result data set.
2. The method for number network search sorting based on data pragmatics according to claim 1, wherein The determining, from the registry of the digital network distributed in multiple scattered data spaces respectively, a retrieval result data set that matches a retrieval text used to describe a target digital object to be retrieved includes: Extract matching text for the target digital object from the retrieval text; Match the matching text with the metadata of the digital objects stored in each data space recorded in the registry of each data space respectively, so as to determine the first digital object from at least one of the data spaces according to the matched target metadata, and determine the second digital object according to the pragmatic association relationship between the digital objects stored in the data space recorded in the registry corresponding to the determined first digital object; Generate the retrieval result data set according to the matched first digital object and the second digital object.
3. The method for number network search sorting based on data pragmatics according to claim 2, wherein The matching the matching text with the metadata of the digital objects stored in each data space recorded in the registry of each data space respectively, so as to determine the first digital object from at least one of the data spaces according to the matched target metadata, and determine the second digital object according to the pragmatic association relationship between the digital objects stored in the data space recorded in the registry corresponding to the determined first digital object includes: Obtain the registry deployed in each data space, and match the matching text with the metadata of the digital objects in the data space recorded in each registry, so as to determine the first digital object from the data space according to the matched target metadata; Determine the second digital object corresponding to the determined first digital object respectively from each registry according to the preset pragmatic depth association relationship.
4. The method for number network search sorting based on data pragmatics according to claim 1, wherein The determining the query similarity of the metadata of each digital object in the retrieval result data set to the retrieval text respectively includes: Extract the retrieval word segmentation set of the retrieval text from the retrieval text, and extract the respective metadata word segmentation sets from the metadata of the digital objects in each retrieval result data set; Determine the metadata word frequency of each retrieval token in each of the metadata token sets, and determine the inverse document frequency of each retrieval token in each of the metadata token sets in all; Determine the weighted sum of the metadata word frequency and the inverse document frequency of each retrieval token in each retrieval result dataset for each digital object as the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text respectively.
5. The data network search and sorting method based on data pragmatics according to claim 1, wherein The determining the pragmatic weight value of each digital object in the retrieval result dataset includes: Establish a Markov chain model for characterizing the pragmatic association relationship between digital objects in the retrieval result dataset; the established Markov chain model is used to define the pragmatic weight value between digital objects with a direct association relationship as the association probability value between the digital objects with a direct association relationship in the Markov chain model, and define the weight numerical relationship between the pragmatic weight values between digital objects without a direct association relationship as the probability numerical relationship between the association probability values between digital objects without a direct association relationship in the Markov chain model; Solve the Markov chain model to obtain the pragmatic weight value of each digital object.
6. The method for sorting search results in a data Internet of Things based on data pragmatics according to claim 5, wherein, The solving the Markov chain model to obtain the pragmatic weight value of each digital object includes: Initialize the pragmatic weight value of each digital object in the retrieval result dataset to a preset initial pragmatic weight value; Update the sum of the product of the ratio of the first pragmatic weight value and the second pragmatic weight value of each third digital object in the retrieval result dataset, the product of the product and the pragmatic weight value of the third digital object, the damping factor coefficient of the retrieval result dataset, and the pragmatic weight evaluation correction value as the pragmatic weight value of each fourth digital object; the third digital object is the digital object on the upper-level node of the fourth digital object on each Markov chain in the Markov chain model; the first pragmatic weight value is the pragmatic weight value from the third digital object to the fourth digital object on the Markov chain where the third digital object is located; the second pragmatic weight value is the sum of the pragmatic weight values of all digital objects other than the third digital object to the fourth digital object; the pragmatic weight evaluation correction value is the probability value of the fourth digital object forming a pragmatic association relationship with any digital object in the retrieval result dataset.
7. The method for sorting and ranking number network search based on data pragmatics according to claim 1, wherein, The determining the retrieval result evaluation value of each digital object in the retrieval result dataset according to the respective query similarity and pragmatic weight value of each digital object in the retrieval result dataset to obtain the sorting result of the digital objects in the retrieval result dataset includes: Determine the weighted sum of the respective query similarity and pragmatic weight value of each digital object in the retrieval result dataset as the retrieval result evaluation value of each digital object in the retrieval result dataset; Arrange the digital objects according to the magnitudes of the retrieval result evaluation values of each digital object in the retrieval result dataset, so as to obtain a sorting result of the digital objects in the retrieval result dataset.
8. A data pragmatics-based number network search and sorting system, characterized in that, Including: A dataset module, configured to determine, from the registries of the digital network respectively distributed in multiple dispersed data spaces, a retrieval result dataset that matches the retrieval text used to describe the target digital object to be retrieved; the retrieval result dataset includes a first digital object that matches the retrieval text, and a second digital object that is pragmatically associated with the first digital object; A parameter calculation module, configured to determine the query similarity of the metadata of each digital object in the retrieval result dataset to the retrieval text respectively, and determine the pragmatic weight value of each digital object in the retrieval result dataset; An evaluation and sorting module, configured to determine the retrieval result evaluation value of each digital object in the retrieval result dataset according to the query similarity and the pragmatic weight value of each digital object in the retrieval result dataset, so as to obtain a sorting result of the digital objects in the retrieval result dataset.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the digital network search and sorting method based on data pragmatics according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, it implements the steps of the digital network search and sorting method based on data pragmatics according to any one of claims 1 to 7.
Citation Information
Patent Citations
Search result sorting method and device, electronic equipment and storage medium
CN114020866A
Question and answer method and device and computer readable storage medium
CN114490957A
Data query method and device, computer equipment and storage medium
CN118394896A
Data retrieval method and device
CN120045521A
Fatigue-reducing compositions containing naturally derived ingredients
KR102633006B1