Information display methods, devices, electronic equipment, storage media and program products
By dynamically and synchronously identifying the context and business information of input information obtained from web page documents, the system matches and displays the identifiers and descriptions of data assets, solving the problem of low information display efficiency in enterprise-level data development and achieving real-time input, real-time display, and efficient recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-02
AI Technical Summary
In enterprise-level data development and analysis, the identifiers of data assets are often lengthy and highly similar, forcing data personnel to switch between different pages to view data assets, which reduces the efficiency of information display.
By acquiring the currently input information, its context, and business information during the information input process of web page documents, and performing dynamic synchronous recognition, matching and displaying the identifiers of data assets and their business descriptions, the system achieves recognition while inputting data.
It improves the efficiency and accuracy of data asset information display, enables instant input and instant display, and reduces the need for page switching.
Smart Images

Figure CN122132640A_ABST
Abstract
Description
Technical Field
[0001] It relates to the field of data processing technology, specifically to information display methods, devices, electronic equipment, storage media, and program products. Background Technology
[0002] In enterprise-level data development and analysis scenarios, data personnel typically need to input or search for data asset identifiers to perform operations such as query writing, task configuration, and report development. However, in complex enterprise data systems, data asset identifiers are often lengthy and highly similar, requiring data personnel to switch between different pages to view the identifiers of data assets under various business domains, thus impacting the efficiency of information display for data assets. Summary of the Invention
[0003] In some cases, an information display method, device, electronic device, storage medium, and program product are provided to address the problem of low efficiency in displaying information corresponding to data assets.
[0004] Firstly, an information display method includes:
[0005] During the information input process for a webpage document, obtain the first piece of information currently being input; Obtain the context information of the first information and the business information of the webpage document; Based on the first information, the context information of the first information, and the business information, dynamic synchronous identification is performed to obtain the identifier of the first data asset that matches the first information; The web page document displays the identifier of the first data asset and a business description of the first data asset.
[0006] Secondly, an information display device includes: The information acquisition module is used to acquire the first piece of information currently being input during the information input process for a web page document; The scene acquisition module is used to acquire the context information of the first information and the business information of the web page document; The dynamic identification module is used to perform dynamic synchronous identification based on the first information, the context information of the first information, and the business information to obtain the identifier of the first data asset that matches the first information. An information display module is used to display the identifier of the first data asset and a business description of the first data asset in the web page document.
[0007] Thirdly, an electronic device includes: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the information display method described in the first aspect or any corresponding embodiment.
[0008] Fourthly, a computer-readable storage medium storing computer instructions for causing a computer to perform the information display method described in the first aspect or any corresponding embodiment thereof.
[0009] Fifthly, a computer program product includes computer instructions for causing a computer to execute the information display method described in the first aspect or any corresponding embodiment thereof.
[0010] In some cases, information display methods involve acquiring the initial information input during the process of inputting information into a webpage document; acquiring the contextual information of the initial information and the business information of the webpage document; dynamically and synchronously identifying the first data asset matching the initial information based on the initial information, its contextual information, and the business information; and displaying the identifier and business description of the first data asset in the webpage document. This method improves recognition efficiency by dynamically and synchronously recognizing the initial information input by the user during input. Furthermore, combining contextual and business information enriches the information used for recognition, improving accuracy. Finally, the identified identifier and business description of the first data asset are displayed in the webpage document, achieving real-time input and real-time display, thus improving the efficiency of information display corresponding to data assets. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the specific embodiments or related technologies, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 These are schematic diagrams based on application scenarios under certain conditions; Figure 2 This is a flowchart illustrating the first method of information display under certain circumstances; Figure 3 These are illustrations based on browser pages under certain conditions; Figure 4 This is a diagram based on the first window in some situations; Figure 5This is a flowchart illustrating the second method of information display under certain circumstances; Figure 6 This is a schematic diagram illustrating the working principle of information display methods under certain circumstances; Figure 7 This is a system architecture diagram based on information display methods under certain circumstances; Figure 8 This is a structural block diagram of an information display device under certain circumstances; Figure 9 These are schematic diagrams of the hardware structure of electronic devices under certain conditions. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages in some cases clearer, the technical solutions in some cases will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are partial embodiments, not all embodiments. Based on the embodiments in some cases, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this solution.
[0014] It is understood that before using the technical solutions disclosed in each embodiment, users should be informed of the type, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and user authorization should be obtained.
[0015] For example, upon receiving a user's proactive request, a prompt message can be sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation based on the prompt message.
[0016] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0017] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the specific implementation method. Other methods that comply with relevant laws and regulations may also be applied to this implementation method.
[0018] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0019] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this document, "multiple" means two or more, unless otherwise explicitly specified.
[0020] When users input information into web documents, they typically desire predictive text input that suggests the identifier of the data asset they want to input without requiring them to type the complete identifier. However, data asset identifiers are often lengthy and complex, making them difficult for users to remember and input accurately. Furthermore, inconsistencies between the vocabulary in business information and the naming conventions of data asset identifiers prevent direct prediction through prefix matching. Because the input fields in web documents cannot understand the semantics or business context of the current document, the predicted data asset identifiers are often irrelevant.
[0021] In related technologies, predictive input methods typically employ prefix matching, static popularity ranking, offline keyword tables, and model recall. Prefix matching generates candidate identifiers based on string prefix matching; however, the naming conventions of words in business information and data assets are inconsistent, causing this method to fail to recognize business semantics.
[0022] Static popularity ranking generates candidate identifiers by sorting them according to global usage frequency or time. Because this method lacks context awareness, the generated results are irrelevant.
[0023] Offline keyword tables require manual maintenance of the mapping relationship between data asset identifiers and keywords, which is costly to maintain and not updated in a timely manner.
[0024] Model recall requires the introduction of simple embedding vectors for retrieval, but this approach does not combine user input with the context of the web page document and cannot provide personalized learning.
[0025] Based on this, in some cases, an information display method is provided, which involves obtaining the first information currently input during the information input process of a web page document; obtaining the context information of the first information and the business information of the web page document; performing dynamic synchronous identification based on the first information, the context information of the first information, and the business information to obtain the identifier of the first data asset that matches the first information; and displaying the identifier of the first data asset and the business description of the first data asset in the web page document.
[0026] This method improves recognition efficiency by dynamically and synchronously recognizing the first information input by the user through input-while-recognition. On this basis, it combines contextual information and business information to enrich the information used for recognition and improve the accuracy of recognition. Finally, the identifier of the first data asset and its business description are displayed in a web page document, realizing real-time input and real-time display and improving the efficiency of information display corresponding to data assets.
[0027] As an optional application scenario in some situations, such as Figure 1 As shown, application 101 is installed in terminal device 110, and user 130 can interact with application 101 through terminal device 110 and / or access device of terminal device 110.
[0028] For example, application 101 could be a browser application. Figure 1 In the application scenario shown, if application 101 is active, the terminal device 110 can display the interface 102 of application 101. The interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.
[0029] In some cases, terminal device 110 communicates with server 120 to provide services to application 101. Terminal device 110 can be a mobile terminal, fixed terminal, or portable terminal, etc., including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of interface, and server 120 can be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0030] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection.
[0031] The following description, in conjunction with the accompanying drawings, illustrates several scenarios. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The various graphic elements on the page may have different arrangements and visual representations; one or more elements may be omitted or replaced, and one or more other elements may also be present, without any limitation in some cases. Furthermore, the embodiments described below primarily pertain to terminal device 110. It should be understood that the actions described relative to terminal device 110 can be performed by application 101 on terminal device 110, or by application 101 in collaboration with its server (e.g., server 120).
[0032] For example, users view web page data in application 101, including but not limited to online documents, data requirement documents, etc. When users input identifiers in web page documents, the system displays matching data asset identifiers and business descriptions through as-entry suggestions, making it easy for users to view.
[0033] It should be understood that, in some cases, the information display method may be provided by installing a plugin in the browser application, enabling the browser application to have the information display function described in some cases; or, it may be by providing a new browser application that integrates the information display function described in some cases.
[0034] In some cases, an embodiment of an information display method is provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0035] In some cases, an information display method is provided that can be used on terminal devices. Figure 2 It is a flowchart based on information display methods in certain situations, such as Figure 2 As shown, the process includes the following steps: Step S201: During the information input process for the webpage document, obtain the first information currently input.
[0036] Web page documents are used to represent documents displayed in browser applications, such as online documents, operation and maintenance panels, configuration interfaces, requirements specifications, etc. Users can edit web page documents by interacting with them.
[0037] Users interact with a webpage document, inputting information in any input area. During this process, the system captures the first piece of information entered. Specifically, by detecting the user's input, subsequent recognition is performed while detecting the first input character.
[0038] For example, a web page document provides an input box. Users interact with the input box to input characters. Correspondingly, the input characters are detected and used as the first piece of information. It should be understood that confirmation of the input information generally requires a confirmation and submission operation, such as pressing Enter, interacting with a confirmation control, etc. The first piece of information currently input is used to characterize the currently input information that has not yet been confirmed and submitted.
[0039] Optionally, for browser plugins, content scripts are injected into the web page document to detect interactive events on the web page document, including but not limited to interactions with rich text areas and custom editing layers of various online documents or ticketing systems.
[0040] Step S202: Obtain the context information of the first information and the business information of the web page document.
[0041] The contextual information of the first piece of information is used to represent natural language segments near the input position. For example, the three lines of description before and after the current input position. Exemplarily, the context of the first piece of information is a short context near the input cursor, such as an entire sentence near the cursor. Obtaining the contextual information of the first piece of information provides contextual information for subsequent semantic understanding.
[0042] The business information in a webpage document is used to characterize the business represented by the webpage document. For example, this includes extracting the page title, Uniform Resource Locator (URL) path, location, and domain terms appearing in the webpage document. Domain term identification can be achieved by retrieving the content of the webpage document and comparing it with a domain database to determine the domain terms appearing in the webpage document.
[0043] The information obtained from the web page document can be represented using a structure to obtain the business information of the web page document. Alternatively, the information obtained from the web page document can be used to classify the web page document into business scenarios to obtain business scenario tags.
[0044] Step S203: Based on the first information, the context information of the first information, and the business information, dynamic synchronous identification is performed to obtain the identifier of the first data asset that matches the first information.
[0045] Dynamic synchronous recognition means that input is detected and recognized simultaneously, rather than waiting for the user to confirm the input before recognition. It should be understood that synchronization does not specifically mean that input and recognition occur strictly simultaneously, but rather allows for a certain degree of systematic error.
[0046] Dynamic synchronous identification can perform parallel matching on the identifiers of optional data assets using the first information, the context information of the first information, and business information. The matching results are then merged and deduplicated to obtain the identifier of the first data asset that matches the first information. Optional data assets can be all data assets within the enterprise, or data assets divided by department. The scope of optional data assets is not limited and can be set according to actual needs.
[0047] For example, the first information can be matched with the summaries of each optional data asset to obtain a matching result; the context information of the first information can be semantically matched with the text descriptions of each optional data asset to obtain a matching result; the business information can be matched with the business information of each optional data asset, such as matching business scenarios, etc., to obtain a matching result. After each matching, the identifier of the matched data asset can be obtained, and the three matching results are then merged and deduplicated to obtain the identifier of the first data asset that matches the first information.
[0048] It should be noted that there may be one or more identifiers for the first data asset. Furthermore, the identifier of the first data asset may also change as the initial information changes; this will be handled according to actual needs, and no restrictions are placed on it here.
[0049] It should be understood that the identifier of a data asset can be related to the type of the data asset. If the data asset is a data table, the identifier of the data asset can be the table name; if the data asset is a text file, the identifier of the data asset can be the file name; and so on.
[0050] Step S204: Display the identifier of the first data asset and its business description in the web page document.
[0051] For each data asset, a language model can be used to generate a corresponding business description, and the identifier and business description of the data asset can be stored in the form of key-value pairs.
[0052] During the user's input of initial information, the system undergoes dynamic synchronization identification in step S203 described above to obtain the identifier of the first data asset matching the initial information. The identifier of the first data asset and its business description are then displayed in the webpage document. If multiple identifiers for the first data asset exist, each identifier and its business description are displayed sequentially.
[0053] The business description of the first data asset, including but not limited to the semantic interpretation, domain affiliation, and timeliness attributes of the first data asset, is used to ensure decision-making. There are no restrictions on its specific content here, and it can be set according to actual needs.
[0054] For example, the business purpose, definition, and upstream and downstream indicators of the data asset are input into the summary generation model to obtain a natural language summary of the data asset, which is then used as the business description of the data asset.
[0055] In some cases, information display methods utilize a input-and-recognition approach, dynamically and synchronously recognizing the initial information input by the user, thus improving recognition efficiency. Building upon this, combining contextual and business information enriches the information used for recognition, enhancing accuracy. Finally, the identified identifier and business description of the first data asset are displayed in a webpage document, achieving instant input and instant display, thereby improving the efficiency of information display corresponding to the data asset. Furthermore, this method, by combining contextual and business information from the initial information, evolves from guessing the desired input string to inferring the specific business scenario for which the data asset is needed.
[0056] In some alternative implementations, step S204 includes: Step a1: Display the identifier of the first data asset and its business description in the first window of the web page document.
[0057] Step a2: If the location of the interaction point is detected to be within the area of the first window, then display the detailed information of the first data asset.
[0058] The web page document displays the identifier of the first data asset and its business description in a first window, for example, such as... Figure 3 As shown, an input box 302 is displayed on the browser page 301. The user inputs information by interacting with the input box 302. Correspondingly, after the input information is detected, it is used as the first information for dynamic synchronous identification to obtain the identifier of the first data asset that matches the first information. The identifiers and business descriptions of each data asset are then displayed sequentially in the first window 303.
[0059] It should be understood that, Figure 3 This is just an example and does not limit the scope of protection. You can set it up according to your actual needs.
[0060] Regarding the identifiers and business descriptions of the first data assets displayed in the first window, due to the limited size of the first window, the business descriptions of each first data asset displayed in the first window are only brief descriptions. Therefore, detailed information about the corresponding data asset can be viewed through the location of the interaction point. The interaction point can be used to indicate the current position of the mouse, and is not limited to the position point corresponding to the interaction operation. That is, the position of the interaction point represents the position where the mouse hovers.
[0061] By detecting the location of the interaction point, if the interaction point is located within the area of the first window, the detailed information of the first data asset is displayed. This detailed information can be the detailed information of all first data assets, or it can be the detailed information of the first data asset pointed to by the interaction point. Furthermore, the detailed information of the first data asset can be displayed in a new window or in other ways; there are no restrictions on this, and it can be set according to actual needs.
[0062] The detailed information of the first data asset may include a description of its purpose, the business domain to which it belongs, and related data assets, etc., but no specific restrictions are placed on its content here.
[0063] The web page document uses a first window to display the identifier and business description of the first data asset. When the interaction point is located in the area of the first window, the details are displayed, which allows you to view the details of the data asset without switching pages.
[0064] In some alternative implementations, if there are more than one identifier for the first data asset, then step a2 above includes: Step a21: Detect the identifier of the second data asset corresponding to the location of the interaction point in the first window.
[0065] Step a22 displays the details of the second data asset.
[0066] For example, such as Figure 4 As shown, in the first window 401, each data asset has a corresponding area for its identifier and business description. Figure 4 The area shown is 402.
[0067] By detecting the region where the interaction point is located and combining this with the mapping relationship between the region and the identifier of the data asset, the identifier of the second data asset to which the interaction point points can be determined. Accordingly, only the detailed information of the second data asset is displayed, achieving on-demand display and highlighting content of interest.
[0068] It should be understood that the identifier of the second data asset is one of the identifiers of multiple first data assets, and not the identifier of a new data asset.
[0069] For cases with more than one identifier for a first data asset, the system displays detailed information about the corresponding data asset based on the position of each identifier in the first window and the position of the interaction point, thus achieving on-demand display.
[0070] In some alternative implementations, the above information display method further includes: Step b1: Obtain the identifier of the target data asset from the identifiers of the first data asset.
[0071] Step b2: Replace the first information with the identifier of the target data asset, and display the identifier of the target data asset in the interactive area of the first information.
[0072] Since the identifiers and business information of each first data asset that matches the first information have been displayed, the user can select one of the identifiers of the first data assets as the identifier of the target data asset as needed. After selection, the identifier of the target data asset replaces the first information and is displayed in the interactive area of the first information.
[0073] For example, such as Figure 4 As shown, if the identifier 1 of the data asset is selected as the identifier of the target data asset, the identifier of the target data asset will be filled into the input area of the web page document to replace the first information.
[0074] The system automatically completes the data asset identification by replacing the currently entered first information with the identifier of the selected target data asset.
[0075] In some cases, an information display method is provided that can be used on terminal devices. Figure 5 It is a flowchart based on information display methods in certain situations, such as Figure 5 As shown, the process includes the following steps: Step S501: During the information input process for the webpage document, obtain the first piece of information currently being input. See details below. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0076] Step S502: Obtain the context information of the first information and the business information of the webpage document. See details... Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0077] Step S503: Based on the first information, the context information of the first information, and the business information, dynamic synchronous identification is performed to obtain the identifier of the first data asset that matches the first information.
[0078] Specifically, using a parallel multi-retrieval strategy, the following three types of recall channels are executed for each input request, and the recall results are then fused to obtain the identifier of the first data asset matching the first information. The three types of recall channels include, but are not limited to, character matching, text relevance matching, and semantic matching. Based on this, step S503 includes: Step S5031: Based on the first information, perform character matching with the identifiers of each optional data asset to obtain the first matching result corresponding to each optional data asset.
[0079] For each optional data asset identifier, the first information is used to perform character matching to obtain literal similarity. For example, literal similarity score can be used to represent the first matching result.
[0080] For example, using the first information and its possible splits, a prefix or substring matching algorithm is performed between the identifiers of each optional data asset and the similarity matching algorithm to obtain the similarity score corresponding to each optional data asset. A threshold can be set to retain only the identifiers of optional data assets that are higher than the threshold as the first matching result. Of course, it is also possible to retain the identifiers of all optional data assets and use the identifiers of all optional data assets and their similarity scores as the first matching result.
[0081] It should be understood that the similarity matching algorithm can be set according to actual needs. For example, an algorithm that can tolerate spelling errors, typos, and incorrect suffixes can be used, or other algorithms can be used. There are no restrictions on it here. The specific algorithm can be set according to actual needs.
[0082] Step S5032: Based on the first information and context information, perform text relevance matching with the text description of each optional data asset to obtain the second matching result corresponding to each optional data asset.
[0083] The text descriptions of each optional data asset include, but are not limited to, the Chinese alias of the optional data asset, the description of its purpose, the description of the indicators / calibers, the description of the output tasks, and the introduction of the business domain, etc., without any specific restrictions on them.
[0084] Using the initial information and contextual information, information retrieval techniques are employed to perform text relevance matching with the text descriptions of each optional data asset, resulting in a relevance score for each optional data asset. Similarly, a relevance threshold can be set, retaining only the identifiers of optional data assets with relevance scores higher than the threshold as the second matching result. Alternatively, all identifiers of optional data assets can be retained, and all identifiers of optional data assets and their relevance scores can be used as the second matching result.
[0085] In the process of text relevance matching, using information retrieval techniques rather than language model matching can reduce the illusions brought about by language models and improve the accuracy of matching results. That is, this matching method is a repeatable engineering process at the information retrieval level, which can be implemented and verified, and can be used for data asset-level association retrieval.
[0086] Step S5033: Based on the first information and the semantic vectors of each optional data asset, perform semantic matching to obtain the third matching result corresponding to each optional data asset.
[0087] For each optional data asset, a corresponding semantic vector is generated in advance. The content of the semantic vector includes, but is not limited to, identifiers, annotations, business domains, and downstream indicator uses. The specific settings are based on actual needs, and no restrictions are imposed on them here.
[0088] By encoding the first piece of information, a corresponding information vector is obtained. This information vector is then semantically matched with the semantic vectors of each optional data asset to obtain a semantic matching result, such as a semantic score. Similarly, a semantic threshold can be set, retaining only the identifiers of optional data assets with semantic scores higher than the threshold as the third matching result. Alternatively, all identifiers of optional data assets can be retained, and all identifiers of optional data assets and their semantic scores can be used as the third matching result.
[0089] In some alternative implementations, the text descriptions of each optional data asset are configured to be generated using a text generation model and the associated information of the data asset; the semantic vectors of each optional data asset are configured to be generated using a vector generation model and the associated information of the data asset.
[0090] In some cases, the relevant content of each optional data asset can be utilized offline using the corresponding generative models. Specifically, a text generation model is used to generate text descriptions of each optional data asset; a vector generation model is used to generate semantic vectors of each optional data asset. Both the text generation model and the vector generation model can be built based on a Large Language Model (LLM).
[0091] For example, the input to the text generation model includes the associated information of optional data assets, including but not limited to business uses, definitions, and descriptions of upstream and downstream indicators, etc., without any limitation on them; the output of the text generation model is a text description.
[0092] The input to the semantic vector generation model includes the association information of optional data assets, including but not limited to identifiers, annotations, business domains, and upstream and downstream indicator uses; the output of the semantic vector generation model includes semantic vectors.
[0093] For each generative model, the solution is not to provide answers online, but to languageize and contextualize data assets and pre-solidify them into structured knowledge for immediate retrieval and sorting.
[0094] In some alternative implementations, step S5033 includes: Step c1: Encode the first information to obtain the first query vector.
[0095] Step c2: Use the first information to query the extended dictionary to obtain the extended information of the first information. The entries in the extended dictionary are configured to be generated using the dictionary generation model and the association information of the data assets.
[0096] Step c3: Encode the extended information to obtain the second query vector.
[0097] Step c4 involves semantic matching of the first query vector and the second query vector with the semantic vectors of each optional data asset to obtain the third matching result corresponding to each optional data asset.
[0098] The expanded dictionary is generated offline using a dictionary generation model. The input to the dictionary generation model includes the business purpose, terminology, and contextual indicators of the data assets. Hint words instruct the dictionary generation model to generate entries for common questions, colloquialisms, business alternatives, and slang terms related to the data assets. These entries are then organized to form the dictionary generation model.
[0099] By using the first information to query the extended dictionary, extended information with semantic similarity to the first information is obtained. The extended information is used as additional query terms to expand the recall coverage and improve the availability during the cold start phase.
[0100] Specifically, the first information and the extended information are encoded to obtain a first query vector and a second query vector, respectively. Then, semantic matching is performed between the first query vector and the semantic vector of each optional data asset, and semantic matching is performed between the second query vector and the semantic vector of each optional data asset, respectively, to obtain the corresponding semantic matching results. The semantic matching results of the two sets of results are merged and deduplicated to obtain the third matching result corresponding to each optional data asset.
[0101] When performing semantic matching, the extended information obtained from the extended dictionary is combined to enrich the query vector; and the extended dictionary is used to represent the data assets into a set of synonyms for multiple colloquial, aliased and business-specific expressions, providing the basic conditions for semantic matching.
[0102] Step S5034: Based on the fusion result of the first matching result, the second matching result, and the third matching result, as well as the business information, the identifiers of the optional data assets are filtered to obtain the identifiers of the first data assets that match the first information.
[0103] The fusion of the first, second, and third matching results can be achieved by setting weights for each of these three matching results, and then weighting these weights with the scores obtained from the corresponding matching results to calculate the fusion score for the optional data assets. It should be understood that if the first to third matching results already contain the identifiers of the filtered optional data assets, then the fusion result of these three matching results may only contain the fusion score for some of the optional data assets, and not all of them.
[0104] Furthermore, business information can be matched with the business information of each optional data asset to obtain a business matching score. This score is then combined with the weight of the business matching and the combined score of the three matching types mentioned above to obtain the total score of the optional data assets. Based on the total score, the top N assets are selected as the identifiers of the first data assets that match the first information. Here, N is a positive integer greater than zero, and the value of N is set according to actual needs; no restrictions are placed on it here.
[0105] In some alternative implementations, step S5034 includes: Step d1: Based on the fusion result of the first matching result, the second matching result, and the third matching result, the identifiers of the optional data assets are filtered to obtain the identifier of the third data asset and the first matching degree corresponding to the identifier of the third data asset.
[0106] Step d2: Obtain the second matching degree between the business domain and business information of each third data asset.
[0107] Step d3: Obtain the personalized scores of the target object for each third data asset to obtain the third matching degree.
[0108] Step d4: Obtain the popularity score of each third data asset to obtain the fourth matching degree.
[0109] Step d5: Based on the weighted fusion results of the first matching degree, the second matching degree, the third matching degree, and the fourth matching degree, the identifiers of each third data asset are filtered to obtain the identifiers of the first data assets that match the first information.
[0110] By fusing the first, second, and third matching results, the identifiers of the selectable data assets are filtered to obtain the identifiers of the third data assets and their first matching scores. The first matching score represents the fusion score of the three matching categories. Each identifier of a third data asset has a corresponding first matching score.
[0111] For each third data asset, the business domain and business information of the third data asset are matched to determine the degree of fit between the two, and a second matching degree is obtained.
[0112] For each third data asset, personalized scoring is used to characterize the target object's attention to the third data asset, or the attention to the business domain to which the third data asset belongs, etc., to obtain the third matching degree.
[0113] For each third-party data asset, the popularity score is used to represent the overall popularity, such as the access frequency and citation count of the third-party data asset, to obtain the fourth matching degree.
[0114] Based on the weighted fusion results of the first matching degree, the second matching degree, the third matching degree, and the fourth matching degree, the score value of each third data asset is obtained, and the identifiers of the N third data assets with the highest scores are used as the identifiers of the first data assets that match the first information.
[0115] The weights of the first, second, third, and fourth matching degrees are set according to actual needs, and no restrictions are imposed on them here. They can be set according to actual needs.
[0116] Based on the three matching methods, personalized scoring and popularity scoring are combined for filtering to select the most suitable identifiers for data assets in the current context, rather than just character matching.
[0117] Step S504: Display the identifier of the first data asset and its business description in the web page document. See details. Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0118] In some cases, information display methods use character matching, text relevance matching, and semantic matching—multi-dimensional matching results—to filter the identifiers of selectable data assets, thereby improving the accuracy and reliability of the filtering results.
[0119] In some alternative implementations, the above information display method further includes: Step e1: Update the personalized score and popularity score of the first data asset.
[0120] Step e2: Update the extended dictionary based on the matching relationship between the first information and the first data asset.
[0121] After identifying the first data asset, the personalized score and popularity score of the first data asset are updated to enable subsequent dynamic synchronization and identification.
[0122] Furthermore, since the identifier of the first data asset is matched with the first information, that is, there is a matching relationship between the first information and the first data asset, the first information can be used as a synonym for the first data asset, and the extended dictionary can be updated.
[0123] The filtering results are used to update the personalized score, popularity score, and expanded dictionary. Through continuous updates, the accuracy of subsequent query matching and personalized intelligent completion are improved.
[0124] As a specific application example in some situations, users need to reference data table names when editing online documents in a browser application. Based on this, such as... Figure 6 As shown, users input characters or phrases through interaction with online documents. The input detection module obtains the current input token, the context extraction module generates a context vector, and the candidate recall module performs parallel recall using multiple retrieval machines to obtain the recall results. The ranking decision module calculates multi-feature scores, and the candidate display module displays the top N (Top N) candidate input tokens, their table names, and business descriptions by rendering floating candidate cards. Users interact with the floating candidate cards to select the target table from the candidates and insert the table name into the original input position. Correspondingly, the personalized learning module updates the weights or popularity.
[0125] Furthermore, such as Figure 7 As shown, the implementation of the information display method is divided into two parts: online recognition and display (701) and offline semantic asset construction (702). The online recognition and display relies on multiple modules: input detection, candidate recall, context extraction, ranking decision, candidate display, interactive feedback and selection, and personalized learning. Offline semantic asset construction is based on data asset metadata, such as table names, annotations, and uses. It utilizes an embedding vector generation model to generate semantic vector indices for the data assets, uses an LLM semantic extension word generation model to generate and cache a large-scale extended dictionary, and uses an LLM use explanation generation model to generate and cache table use explanations. The LLM-generated semantic extension words and use explanations are pre-computed and cached offline; online, only index retrieval and weighted fusion are used to achieve real-time response.
[0126] This method employs a multi-retrieval parallel recall mechanism, introducing three heterogeneous retrieval algorithms to be executed in parallel within the browser completion system. It comprehensively captures literal, semantic, and textual table name associations to form a high-recall candidate pool. Through a large language model, offline semantic distillation is performed on the metadata, field descriptions, and usage text of enterprise data assets, automatically generating extended word sets such as business aliases, colloquial expressions, and industry terms to enhance semantic recall coverage and support cross-language and cross-domain completion.
[0127] This method is also a weighted ranking algorithm that integrates multiple features such as input similarity, text retrieval score, semantic similarity, page business domain, and popularity, so that the recommendation results have contextual relevance and organizational consistency.
[0128] As described above, this method utilizes a browser plugin to implement input detection, candidate display, and interactive feedback in any webpage input box without modifying the host system code, thus building a cross-system deployable intelligent input layer. Furthermore, during user input, differential triggering is performed on changes in input characters, and candidate results are only partially rearranged for newly added characters, significantly reducing computational burden and ensuring real-time performance.
[0129] In some cases, an information display device is also provided to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0130] Information display devices in some situations, such as Figure 8 As shown, it includes: The information acquisition module 801 is used to acquire the first information currently being input during the information input process for a web page document.
[0131] The scene acquisition module 802 is used to acquire the context information of the first information and the business information of the web page document.
[0132] The dynamic identification module 803 is used to perform dynamic synchronous identification based on the first information, the context information of the first information, and business information to obtain the identifier of the first data asset that matches the first information.
[0133] The information display module 804 is used to display the identifier of the first data asset and the business description of the first data asset in a web page document.
[0134] In some alternative implementations, the information display module 804 includes: The first display unit is used to display the identifier of the first data asset and its business description in a web page document using a first window.
[0135] The second display unit is used to display detailed information about the first data asset if the location of the interaction point is detected to be within the area of the first window.
[0136] In some optional implementations, if there are more than one identifier for the first data asset, the second display unit includes: The identifier detection subunit is used to detect the identifier of the second data asset corresponding to the location of the interaction point in the first window.
[0137] The information display sub-unit is used to display detailed information about the second data asset.
[0138] In some alternative implementations, the information display device further includes: The identifier acquisition module is used to acquire the identifier of the target data asset from the identifiers of the first data asset.
[0139] The identifier display module is used to replace the first information with the identifier of the target data asset and display the identifier of the target data asset in the interaction area of the first information.
[0140] In some alternative implementations, the dynamic recognition module 803 includes: The character matching unit is used to perform character matching based on the first information and the identifier of each optional data asset to obtain the first matching result corresponding to each optional data asset.
[0141] The relevance matching unit is used to perform text relevance matching with the text description of each optional data asset based on the first information and context information, so as to obtain the second matching result corresponding to each optional data asset.
[0142] The semantic matching unit is used to perform semantic matching based on the first information and the semantic vectors of each optional data asset to obtain the third matching result corresponding to each optional data asset.
[0143] The filtering unit is used to filter the identifiers of optional data assets based on the fusion results of the first matching result, the second matching result, and the third matching result, as well as business information, to obtain the identifier of the first data asset that matches the first information.
[0144] In some optional implementations, the semantic matching unit includes: The first encoding subunit is used to encode the first information to obtain the first query vector.
[0145] The dictionary query subunit is used to query the extended dictionary using the first information to obtain the extended information of the first information. The entries in the extended dictionary are configured to be generated using the dictionary generation model and the association information of the data assets.
[0146] The second encoding subunit is used to encode the extended information to obtain the second query vector.
[0147] The semantic matching subunit is used to perform semantic matching between the first query vector and the second query vector and the semantic vector of each optional data asset, respectively, to obtain the third matching result corresponding to each optional data asset.
[0148] In some alternative implementations, the text descriptions of each optional data asset are configured to be generated using a text generation model and the associated information of the data asset; the semantic vectors of each optional data asset are configured to be generated using a vector generation model and the associated information of the data asset.
[0149] In some alternative implementations, the filtering unit includes: The first matching subunit is used to filter the identifiers of optional data assets based on the fusion result of the first matching result, the second matching result, and the third matching result, so as to obtain the identifier of the third data asset and the first matching degree corresponding to the identifier of the third data asset.
[0150] The second matching subunit is used to obtain the second matching degree between the business domain and business information of each third data asset.
[0151] The third matching subunit is used to obtain the target object's personalized score for each third data asset, thus obtaining the third matching degree.
[0152] The fourth matching subunit is used to obtain the popularity score of each third data asset to obtain the fourth matching degree.
[0153] The filtering subunit is used to filter the identifiers of each third data asset based on the weighted fusion results of the first matching degree, the second matching degree, the third matching degree, and the fourth matching degree, so as to obtain the identifiers of the first data assets that match the first information.
[0154] In some alternative implementations, the information display device further includes: The first update module is used to update the personalized rating and popularity rating of the first data asset; The second update module is used to update the extended dictionary based on the matching relationship between the first information and the first data asset.
[0155] In some cases, the provided information display device can execute the above-described information display method, possessing the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0156] Figure 9 This is a schematic diagram of the structure of an electronic device provided in certain situations.
[0157] The following is a detailed reference. Figure 9This diagram illustrates a suitable structure for implementing an electronic device in various scenarios. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 901, which performs various appropriate actions and processes based on a program stored in read-only memory (ROM) 902 or loaded from memory 908 into random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of the electronic device. The processor 901, ROM 902, and RAM 903 are interconnected via bus 904. An input / output interface 905 is also connected to bus 904.
[0158] Typically, the following devices can be connected to the input / output interface 905: input devices 906 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 907 including, for example, a liquid crystal display, speaker, vibrator, etc.; memory devices 908 including, for example, magnetic tape, hard disk, etc.; and communication devices 909. Communication devices 909 allow electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0159] Specifically, the processes described in the flowchart above can be implemented as computer software programs. For example, some cases include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication device 909, or installed from memory 908, or installed from ROM 902. When the computer program is executed by processor 901, it performs the functions defined in the information display methods in some cases.
[0160] Figure 9 The electronic devices shown are merely examples and should not be construed as limiting their functionality or scope of use in any situation.
[0161] In some cases, a computer-readable storage medium is also provided, in which the above-described methods can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the information display method shown in the above embodiments.
[0162] Some of the above solutions can be applied as computer program products, such as computer program instructions. When executed by a computer, these instructions, through the operation of the computer, can invoke or provide the aforementioned methods and / or technical solutions. Those skilled in the art should understand that the forms in which computer program instructions exist in computer-readable media include, but are not limited to, source files, executable files, and installation package files. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instruction; the computer compiling the instruction and then executing the corresponding compiled program; the computer reading and executing the instruction; or the computer reading and installing the instruction and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0163] Although embodiments in some cases have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the above description, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. An information display method, comprising: During the information input process for a webpage document, obtain the first piece of information currently being input; Obtain the context information of the first information and the business information of the webpage document; Based on the first information, the context information of the first information, and the business information, dynamic synchronous identification is performed to obtain the identifier of the first data asset that matches the first information; The web page document displays the identifier of the first data asset and a business description of the first data asset.
2. The method according to claim 1, wherein displaying the identifier of the first data asset and the business description of the first data asset in the web page document includes: The first window in the web page document displays the identifier of the first data asset and its business description. If the location of the interaction point is detected to be within the area of the first window, then the detailed information of the first data asset is displayed.
3. The method according to claim 2, wherein if there are more than one identifier for a first data asset, then if the location of the interaction point is detected to be within the area of the first window, the detailed information of the first data asset is displayed, including: The identifier of the second data asset corresponding to the location of the interaction point in the first window is detected; Displays detailed information about the second data asset.
4. The method according to claim 1, further comprising: Obtain the identifier of the target data asset from the identifier of the first data asset; The first information is replaced with the identifier of the target data asset, and the identifier of the target data asset is displayed in the interactive area of the first information.
5. The method according to claim 1, wherein the step of dynamically synchronizing and identifying based on the first information, the context information of the first information, and the business information to obtain the identifier of the first data asset matching the first information includes: Based on the first information and the identifiers of each optional data asset, character matching is performed to obtain the first matching result corresponding to each optional data asset; Based on the first information and the context information, text relevance matching is performed with the text descriptions of each of the optional data assets to obtain a second matching result corresponding to each of the optional data assets; Based on the first information and the semantic vectors of each of the optional data assets, a semantic matching is performed to obtain a third matching result corresponding to each of the optional data assets; Based on the fusion result of the first matching result, the second matching result, and the third matching result, as well as the business information, the identifiers of the optional data assets are filtered to obtain the identifier of the first data asset that matches the first information.
6. The method according to claim 5, wherein the step of semantically matching the first information with the semantic vectors of each of the optional data assets to obtain a third matching result corresponding to each of the optional data assets includes: The first information is encoded to obtain the first query vector; The first information is used to query the extended dictionary to obtain the extended information of the first information. The entries in the extended dictionary are configured to be generated using the dictionary generation model and the association information of the data assets. The extended information is encoded to obtain the second query vector; Semantic matching is performed between the first query vector and the second query vector and the semantic vector of each of the optional data assets to obtain a third matching result corresponding to each of the optional data assets.
7. The method of claim 5, wherein the text description of each of the optional data assets is configured to be generated using a text generation model and the associated information of the data assets; The semantic vectors of each of the optional data assets are configured to be generated using a vector generation model and the associated information of the data assets.
8. The method according to claim 5, wherein the step of filtering the identifiers of the optional data assets based on the fusion result of the first matching result, the second matching result, and the third matching result, and the business information, to obtain the identifier of the first data asset matching the first information, includes: Based on the fusion result of the first matching result, the second matching result, and the third matching result, the identifiers of the optional data assets are filtered to obtain the identifier of the third data asset and the first matching degree corresponding to the identifier of the third data asset; Obtain the second matching degree between the business domain of each of the third data assets and the business information; Obtain the personalized scores of the target object for each of the aforementioned third data assets to obtain the third matching degree; Obtain the popularity score of each of the aforementioned third data assets to obtain the fourth matching degree; Based on the weighted fusion result of the first matching degree, the second matching degree, the third matching degree, and the fourth matching degree, the identifiers of each of the third data assets are filtered to obtain the identifiers of the first data assets that match the first information.
9. The method according to claim 8, further comprising: Update the personalized score and popularity score of the first data asset; Based on the matching relationship between the first information and the first data asset, the extended dictionary is updated.
10. An information display device, comprising: The information acquisition module is used to acquire the first piece of information currently being input during the information input process for a web page document; The scene acquisition module is used to acquire the context information of the first information and the business information of the web page document; The dynamic identification module is used to perform dynamic synchronous identification based on the first information, the context information of the first information, and the business information to obtain the identifier of the first data asset that matches the first information. An information display module is used to display the identifier of the first data asset and a business description of the first data asset in the web page document.
11. An electronic device, comprising: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the information display method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer instructions for causing a computer to perform the information display method according to any one of claims 1 to 9.
13. A computer program product comprising computer instructions for causing a computer to perform the information display method according to any one of claims 1 to 9.