Method and system for classifying one or more hyperlinks in document
By analyzing and classifying text strings in the document, the problem of poor readability of topics when users manage access between multiple related hyperlinks and the main document is solved, and more effective information management and access is achieved.
Patent Information
- Application Number
- CN202380072521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for users to maintain readability of topics in the main document when managing access between multiple related hyperlinks and the main document.
By analyzing text strings in a document, hyperlinks are identified and classified. The method includes analyzing the surrounding text string surrounding the hyperlink and classifying the hyperlink into a predetermined category based on the analysis results.
Improves the readability of multiple hyperlinks in the document, helping users manage and access relevant information more effectively.
Smart Images

Figure CN120077371A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and system for classifying one or more hyperlinks in a document. Background Art
[0002] In the past few years, the huge growth of the web (World Wide Web Internet service) has provided users with a vast amount of information. This information is available in different types of documents. A number of concepts are embedded in such documents in the form of hyperlinks, and in order to better understand these concepts, users need to access many such hyperlinks and ultimately return to the main document. The main problem faced by users is to maintain the readability of the topic in the main document by managing the access between multiple related hyperlinks and the main document. Therefore, there seems to be a need for a solution to increase the readability of the topic in a main document containing multiple hyperlinks.
[0003] The above information is presented only as background information to assist in understanding the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above is applicable prior art with respect to the present disclosure. Summary of the Invention
[0004] Aspects of the present disclosure are to at least address the above problems and / or disadvantages and to at least provide the advantages described below. Accordingly, an aspect of the present disclosure is to provide a method and system for classifying one or more hyperlinks in a document.
[0005] Additional aspects will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the presented embodiments.
[0006] According to an aspect of the present disclosure, there is provided a method for classifying one or more hyperlinks in a document. The method includes identifying one or more hyperlinks in the document based on an analysis of text strings in the document. The method further includes analyzing the surrounding text strings around each of the one or more hyperlinks, and classifying the one or more hyperlinks into at least one of a plurality of predetermined categories based on the analysis of the surrounding text strings around each of the one or more hyperlinks.
[0007] According to another aspect of the present disclosure, there is provided a system for classifying one or more hyperlinks in a document. The system includes an identification unit configured to identify one or more hyperlinks in the document based on an analysis of text strings in the document. The system further includes an analysis unit and a classification unit. The analysis unit is configured to analyze the surrounding text strings around each of the one or more hyperlinks, and the classification unit is configured to classify the one or more hyperlinks into at least one of a plurality of predetermined categories based on the analysis of the surrounding text strings around each of the one or more hyperlinks.
[0008] A detailed description of various embodiments of the present disclosure is disclosed below in conjunction with the accompanying drawings. Other aspects, advantages, and significant features of the present disclosure will become apparent to those skilled in the art. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings, in which:
[0010] Figure 1 shows a method for describing a method for classifying one or more hyperlinks in a document according to an embodiment of the present disclosure;
[0011] Figure 2 shows a block diagram of a system for classifying one or more hyperlinks in a document according to an embodiment of the present disclosure;
[0012] Figure 3 shows various stages of classifying one or more hyperlinks in a document according to an embodiment of the present disclosure;
[0013] Figure 4 shows a block diagram representing the classification of one or more hyperlinks in a document according to an embodiment of the present disclosure;
[0014] Figure 5 shows a link representation according to an embodiment of the present disclosure;
[0015] Figure 6 respectively show an encoder and a decoder of a link representer according to various embodiments of the present disclosure;
[0016] Figure 7A shows various stages of a list of link representations according to an embodiment of the present disclosure;
[0017] Figure 7B shows a multi-head self-attention layer of an encoder and a decoder according to an embodiment of the present disclosure;
[0018] Figure 8Shows a block diagram illustrating the creation of a user knowledge graph according to an embodiment of the present disclosure;
[0019] Figure 9 Shows an example of the creation of a user knowledge graph according to an embodiment of the present disclosure;
[0020] Figure 10 Shows a flowchart for selecting hyperlinks based on user information according to an embodiment of the present disclosure;
[0021] Figure 11A Show the creation of a taxonomy and the taxonomy according to various embodiments of the present disclosure, respectively;
[0022] Figure 11B Show the creation of a taxonomy and the taxonomy according to various embodiments of the present disclosure, respectively;
[0023] Figure 12A Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure;
[0024] Figure 12B Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure;
[0025] Figure 12C Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure;
[0026] Figure 12D Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure;
[0027] Figure 13 Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure;
[0028] Figure 14 Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure;
[0029] Figure 15 Shows various use cases for classifying one or more hyperlinks in a document according to various embodiments of the present disclosure; and
[0030] Figure 16 Shows a comparison of classifying one or more hyperlinks in a document between the prior art and the present disclosure according to an embodiment of the present disclosure.
[0031] Throughout the drawings, it should be noted that the same reference numerals are used to describe the same or similar elements, features, and structures. Detailed Implementation Modes
[0032] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of various embodiments of the present disclosure defined by the claims and their equivalents. It includes various specific details to aid understanding, but these details are only regarded as exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0033] The terms and words used in the following description and claims are not limited to their literal meanings, but are used by the inventor only to enable a clear and consistent understanding of the present disclosure. Thus, it will be apparent to those skilled in the art that the following description of the various embodiments of the present disclosure is for illustrative purposes only and not for the purpose of limiting the present disclosure defined by the appended claims and their equivalents.
[0034] It should be understood that, unless the context clearly indicates otherwise, the singular forms include plural referents. Thus, for example, a reference to "a component surface" includes a reference to one or more such surfaces.
[0035] As used herein, the term "some" is defined as "none, or one, or more than one, or all". Thus, the terms "none", "one", "more than one", "more than one but not all", or "all" will fall within the definition of "some". The term "some embodiments" may refer to no embodiments, one embodiment, several embodiments, or all embodiments. Thus, the term "some embodiments" is defined to mean "no embodiments, or one embodiment, or more than one embodiment, or all embodiments".
[0036] The terms and structures employed herein are used to describe, teach, and illustrate some embodiments and their specific features and elements, and do not limit, restrict, or reduce the spirit and scope of the claims or their equivalents.
[0037] More specifically, unless otherwise specified, any term used herein (such as but not limited to "comprising", "including", "having", "consisting" and their grammatical variants) does not specify an exact limitation or constraint, and does not exclude the possibility of adding one or more features or elements. In addition, unless otherwise stated with the restrictive language "must comprise" or "needs to comprise", the possibility of removing one or more of the listed features and elements shall not be regarded as excluded.
[0038] Regardless of whether a particular feature or element is limited to only one use, in either case it can still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element." Furthermore, unless otherwise specified by restrictive language (such as "requires the presence of one or more ..." or "requires one or more elements"), the use of the term "one or more" or "at least one" feature or element does not exclude the absence of that feature or element.
[0039] Unless otherwise defined, all terms used herein (especially any technical terms and / or scientific terms) may be considered to have the same meaning as commonly understood by one of ordinary skill in the art.
[0040] The present disclosure aims to intelligently create personalized classifications of hyperlinks (i.e., preliminary categories, simultaneous categories, and post-categories) using link indicators when a user browses any document (e.g., web pages, articles, etc.) in real time without accessing the hyperlinks.
[0041] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0042] Figure 1 A method 100 is shown describing a method for classifying one or more hyperlinks in a document according to an embodiment of the present disclosure.
[0043] Figure 2 A block diagram 200 of a system for classifying one or more hyperlinks in a document according to an embodiment of the present disclosure is shown.
[0044] Figure 3 1 shows various stages of classifying one or more hyperlinks in a document according to an embodiment of the present disclosure. Figure 1 , Figure 2 and Figure 3 The descriptions are explained in conjunction with each other.
[0045] Reference Figures 1 to 3 , system 200 may include, but is not limited to, processor 202, memory 204, unit 206, and data unit 208. Unit 206 and memory 204 may be coupled to processor 202.
[0046] Processor 202 may be a single processing unit or several units, all of which may include multiple computing units. Processor 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. Processor 202 is configured to retrieve and execute computer-readable instructions and data stored in memory 204, among other capabilities.
[0047] The memory 204 may include any non - transitory computer - readable medium known in the art, including volatile memory (such as static random - access memory (SRAM) and dynamic random - access memory (DRAM)) and / or non - volatile memory (such as read - only memory (ROM), erasable programmable ROM, flash memory, hard disk, optical disk, and magnetic tape).
[0048] In addition to other components, the unit 206 includes routines, programs, objects, components, data structures, etc. that perform specific tasks or implement data types. The unit 206 can also be implemented as a signal processor, state machine, logic circuit, and / or any other device or component that manipulates signals based on operation instructions.
[0049] The unit 206 can be implemented by hardware, instructions run by a processing unit, or a combination thereof. The processing unit can include a computer, a processor (such as processor 202), a state machine, a logic array, or any other suitable device capable of processing instructions. The processing unit can be a general - purpose processor that runs instructions to cause the general - purpose processor to perform the required tasks, or the processing unit can be dedicated to performing the required functions. In another embodiment of the present disclosure, the unit 206 can be machine - readable instructions (software) that, when run by a processor / processing unit, perform any of the described functions.
[0050] In an embodiment, the unit 206 can include an identification unit 210, an analysis unit 212, a classification unit 214, an extraction unit 216, a selection unit 218, a generation unit 220, and a display unit 222.
[0051] The various units 210 - 222 can communicate with each other. In an embodiment, the various units 210 - 222 can be part of the processor 202. In another embodiment, the processor 202 can be configured to perform the functions of the units 210 - 222. Among other things, the data unit 208 serves as a repository for storing data processed, received, and generated by one or more of the units 206.
[0052] According to an embodiment of the present disclosure, the system 200 can be part of an electronic device on which a document is accessed. According to another embodiment, the system 200 can be incorporated into an electronic device on which a document is accessed. It should be noted that the term "electronic device" refers to any electronic device used by a user, such as a mobile device, a desktop computer, a laptop computer, a personal digital assistant (PDA), or a similar device.
[0053] Refer to Figure 1, at operation 101, the method includes identifying one or more hyperlinks in a document based on an analysis of text strings in the document. In an embodiment, the document may correspond to one of a text document, a web document, an application-based document, a cloud-based document, or any other document.
[0054] Referring Figure 2 and Figure 3 , the identification unit 210 may identify one or more hyperlinks 301a - 301f in the document 303 based on an analysis of text strings in the document 303. For example, by analyzing the text string "Long Short-Term Memory (LSTM) is an artificial neural network used in the fields of artificial intelligence and deep learning", the identification unit 210 may identify hyperlinks 301a - 301c, namely, artificial neural network, artificial intelligence, and deep learning. Similarly, by analyzing the text string "The problem with using gradient descent for standard RNNs is that the error gradient vanishes exponentially fast with the magnitude of the time delay between significant events", the identification unit 210 may identify hyperlinks 301d - 301e, namely, gradient descent and vanishing. Similarly, by analyzing the text string "Although LSTM is considered one of the greatest inventions in NLP, recent models (such as the Transformer) are now more popular", the identification unit 210 may identify hyperlink 301f (Transformer). It should be noted that the identification unit 210 may use techniques known to those skilled in the art to identify one or more hyperlinks.
[0055] Thereafter, at operation 103, method 100 includes analyzing the surrounding text strings around each of the one or more hyperlinks. In an embodiment, the analysis unit 212 may analyze the surrounding text strings around each hyperlink to determine the position of the words in the surrounding text strings relative to the hyperlink. For example, referring Figure 3 , for the surrounding text string around the hyperlink "artificial neural network", the position of each word in the text string may be determined by analyzing the text string, such as the position of the word "LSTM" is the third position to the left of the hyperlink "artificial neural network".
[0056] At operation 105, method 100 may include classifying one or more hyperlinks into at least one of a plurality of predetermined categories based on an analysis of the surrounding text strings around each of the one or more hyperlinks. In an embodiment, the plurality of predetermined categories may include a preparatory category, a simultaneous category, or a post category. The preparatory category refers to a category that recommends that the user access the hyperlink before reading the content of document 303. The simultaneous category refers to a category that recommends that the user access the hyperlink together while reading the content of document 303. The post category refers to a category that recommends that the user access the hyperlink after reading the content of document 303. In an embodiment, classification unit 214 may classify one or more hyperlinks using the surrounding text strings without referring to the one or more hyperlinks. For example, referring to Figure 3 , classification unit 214 may classify each of hyperlinks 301a - 301f into one of the categories. For example, hyperlinks 301a - 301c may be classified into the preparatory category, hyperlinks 301d - 301e may be classified into the simultaneous category, and hyperlink 301f may be classified into the post category, as shown in block 305.
[0057] Referring now to Figure 4 , Figure 5 , Figure 6 , Figure 7A , Figure 7B , Figures 8 to 10 , Figure 11A and Figure 11B the units 307 - 317 shown in Figure 3 are described.
[0058] Figure 4 FIG. shows a block diagram representing the classification of one or more hyperlinks in document 303 according to an embodiment of the present disclosure.
[0059] Referring to Figure 4 , in an embodiment, a BERT model 401 is used to classify one or more hyperlinks in document 303. At block 403, system 200 performs token embeddings on the input text strings in document 303. In an embodiment, WordPiece embeddings are used for token embeddings of the given input text strings. Token embeddings transform words into vector representations of a fixed dimension of 768. Let's consider the following example "LSTM networks are well-suited toclassifying". System 200 first tokenizes the text string into: "lstm net##works are well##-##suited to classify##ing". There are 10 tokens. Additionally, system 200 represents each token as a 768-dimensional vector. Thus, a matrix of shape (10 * 768) is obtained.
[0060] At block 405, system 200 performs link embedding. In an embodiment, two learned embeddings (EL and ENL) of size 768 are used to distinguish link words and non - link words respectively. In an embodiment, link words may refer to words present in a hyperlink, and non - link words may refer to words present in the text string surrounding the hyperlink. For example, ENL is used to represent the first 8 tokens of the link, and EL is used to represent the last two tokens of the link. These embeddings are added element - wise to the token embeddings. Thereafter, system 200 obtains a matrix of shape (10, 768).
[0061] At block 407, system 200 performs positional embedding. In an embodiment, positional embedding is used to feed the position of each word in the text string into the model. The positional embedding is a vector of size 768, which is different for each position. In the example, the positional embeddings are different for all 10 tokens. These embeddings are added element - wise to the matrix obtained at block 407, and finally a matrix of shape (10, 768) is obtained, which is fed into the BERT model 401.
[0062] The final embedding of the classification token (CLS) (which is a 768 - dimensional vector) is obtained from the BERT model 401. The CLS is fed into a hidden neural layer with a weight matrix of size (768, 768). The hidden neural layer provides a new vector of size 768, which is fed into a Softmax layer with a weight matrix of size (3, 768). The Softmax layer provides a 3 - dimensional vector. The Softmax function is applied on the 3 - dimensional vector to obtain the probability of each of the multiple predefined classes.
[0063] One of the predefined classes is assigned to the hyperlink with the maximum probability. For example, referring to Figure 3 , the maximum probability is obtained for the post - category, and thus, the post - category is assigned to the hyperlink "classifying".
[0064] It should be noted that Figure 4 illustrates an embodiment of the present disclosure, and any other suitable model / block / layer can be used to classify hyperlinks.
[0065] In an embodiment, the classified link can be represented by a modified link, where the modified link defines the hyperlink in a more relevant way.
[0066] Figure 5 illustrates a link representer according to an embodiment of the present disclosure. In an embodiment, the link representer 500 may refer to Figure 3 the link representer 307 of
[0067] Referring to Figure 5, the link representer 500 may include a bidirectional and self-attention block 501, "N" encoders 503, "N" decoders 505, and a "unidirectional, autoregressive, and masked self-attention + cross-attention block" 507.
[0068] Figure 6 The encoders and decoders of the link representer 500 according to various embodiments of the present disclosure are shown respectively. They are explained in combination with each other below Figure 5 , Figure 6 .
[0069] Referring to Figure 5 , Figure 6 , first, the input text string of the encoder 503 is tokenized into individual words / sub-words. Subsequently, word embeddings are used to represent the tokenized words. Then, positional encodings are added to each word to preserve the position information of each word in the sentence. As Figure 7A shown, the final input representation of the module is denoted as X. Then, the input X is passed to the encoder 503.
[0070] Figure 7A The various stages of a link representation list according to an embodiment of the present disclosure are shown. Figure 7B The multi-head self-attention layer of the encoder 503 according to an embodiment of the present disclosure is shown.
[0071] Referring to Figure 6 , the encoder 503 includes a multi-head self-attention layer, a normalization layer ( Figure 6 not shown in
[0072] Referring to Figure 6 , the decoder includes a multi-head self-attention layer, an encoder-decoder attention layer, and a feed-forward layer. Figure 7BShows the multi - head self - attention layer of the decoder 505 according to an embodiment of the present disclosure. During training, the input to the decoder 505 is the right - shifted target sentence tokens. In addition, a link embedding is added at each input position to include link information. A masked self - attention layer is used in the decoder 505 such that the prediction at a given token depends only on past tokens. The encoder - decoder attention layer uses the decoder output as the query and the encoder output as the key and value. Thus, the decoder 505 can focus on relevant information from the encoded text. The final decoder representation is fed into an output Softmax layer, which generates a probability distribution for the next output word. At test time, greedy decoding or beam search is used on the word probability distribution generated from the decoder 505 to obtain the final sentence. For example, referring to Figure 3 , as shown in box 309, the output of the link representer 307 for the hyperlink "vanish" is "vanish, vanishing gradient problem".
[0073] In an embodiment, a user may select one or more hyperlinks from the classified hyperlinks. In an embodiment, the extraction unit 216 may extract user information from the memory 204, which stores at least one of the user's user profile information or browsing history information. In an embodiment, the user information may include at least one of the user's profile information, the user's demographic information, the user's education level, published documents written by the user, documents uploaded by the user, and any other information related to the user. In an embodiment, the memory 204 may include a knowledge graph 311 of the user containing the user information. Referring to Figure 8 to explain the creation of the user's knowledge graph 311. Further, the selection unit 218 may select one or more hyperlinks based on the user information associated with the user accessing the document 303.
[0074] Figure 8 Shows a block diagram illustrating the creation of a user knowledge graph according to an embodiment of the present disclosure.
[0075] Referring to Figure 8 , at block 801, the system 200 extracts entities (such as entities and attributes) from the user's structured, unstructured, and / or semi - structured data. The data may include at least one of the user's profile information, the user's demographic information, the user's education level, published documents written by the user, documents uploaded by the user, and any other information related to the user. Thereafter, at block 803, the system 200 links the extracted entities to any existing entities in the knowledge graph 809. However, if the extracted entity does not exist in the knowledge graph 809, the system 200 creates a new entity and adds a triple of the entity (e.g., thermodynamics, property, second law) to the knowledge graph. At block 805, the system 200 applies ontology learning to create the knowledge graph 809. The knowledge graph 809 may be similar toFigure 3 The knowledge graph 311 of the user represented therein. In an embodiment, the system 200 can create a knowledge graph 809 via a data schema 807.
[0076] Figure 9 Shows an example of the creation of a user knowledge graph according to an embodiment of the present disclosure.
[0077] Referring to Figure 9 , the system 200 extracts entities and attributes from semi-structured data as thermodynamics and the second law respectively. Then, the system 200 determines the entity link to be linked to a previously available entity: thermodynamics. Thereafter, the system 200 applies ontology learning to create a user knowledge graph that may include user information (such as name, graduated major, institution, etc.). In another embodiment, in order to select one or more hyperlinks based on user information, the selection unit 218 may include a link selector 315a that selects links based on user information. In an embodiment, the selection unit 218 may also include a hierarchy builder 315b that selects links based on the hierarchy of concepts / topics / user information.
[0078] Figure 10 Shows a flowchart for selecting hyperlinks based on user information according to an embodiment of the present disclosure.
[0079] Referring to Figure 10 , in operation 1001, the system 200 can determine whether a classified hyperlink exists in the user knowledge graph 809. If so, the system 200 can lower the priority of the link because the user has prior knowledge of the link. If not, then in operation 1003, the system 200 can expand the topic of the hyperlink. Then, in operation 1005, the system 200 can determine whether the expanded topic exists in the user knowledge graph. If so, the system 200 can lower the priority of the link because the user has prior knowledge of the topic. If not, the system 200 can select the link to be displayed to the user.
[0080] The generation unit 220 can generate a list of link representations corresponding to one or more classified hyperlinks based on the surrounding text without using the content of one or more hyperlinks. For example, referring to FIG. 7, the list of link representations can be the list shown in FIG. 7.
[0081] After generating the list of link representations, the display unit 222 can display the list of link representations on a graphical user interface (GUI). In an embodiment, the generation unit 220 can create a taxonomy including a list of concepts related to the text string of the document 303.
[0082] Figure 11A And Figure 11B Show the creation of a taxonomy and the taxonomy according to various embodiments of the present disclosure respectively. InFigure 11A The taxonomy created in Figure 3 can refer to the box 313 of
[0083] Refer to Figure 11A , the system 200 can use various extraction techniques (such as Hearst pattern-based candidate extraction, lexicon-based candidate extraction, neural network-based hypernym discovery) for documents and Wikidata (i.e., general data related to the concepts of the documents) to extract relevant data related to the text strings around the hyperlinks. The system 200 can filter the extracted data based on statistical and language patterns and create a taxonomy.
[0084] Refer to Figure 11B , the taxonomy of the concept "artificial intelligence" can include machine learning, deep learning, perceptrons, artificial neural networks.
[0085] The generation unit 220 can arrange one or more classified hyperlinks in a predefined order based on the taxonomy. The display unit 222 can display a list of link representations based on the arrangement of one or more classified hyperlinks in a predefined order. In an embodiment, the display unit 222 can display a list of link representations on the graphical user interface (GUI) of the electronic device. For example, refer to Figure 3 , the display unit 222 can display the list of link representations 317.
[0086] Figure 12A , Figure 12B , Figure 12C , Figure 12D , Figure 13 , Figure 14 and Figure 15 show various use cases of classifying one or more hyperlinks in the document 303 according to various embodiments of the present disclosure.
[0087] Refer to Figure 12A , Figure 12B , Figure 12C and Figure 12D , the user wants to study BERT. According to the disclosed technology, the hyperlinks in the document 303 are classified and displayed as follows:
[0088] Preliminary category:
[0089] Natural Language Processing (NLP)
[0090] Language model
[0091] ULMFiT
[0092] Simultaneous category:
[0093] LSTMs
[0094] Transfer learning in NLP
[0095] Post - category:
[0096] OpenAI's GPT
[0097] Refer to Figure 13 , the user asks the intelligent assistant "How did World War II end?". Then, the intelligent assistant provides some answers. After that, the user asks the intelligent assistant to read the whole article. According to the disclosed technology, the intelligent assistant reads an article with information having hyperlinks classified in the preparatory category, simultaneous category, and post - category.
[0098] Refer to Figure 14 , the home center actively suggests to the user to bake a cake and provides an appropriate preparatory category.
[0099] Refer to Figure 15 , by providing all relevant information in one place before purchasing a product, the user experience is enhanced when shopping on the website / app as follows:
[0100] Preparatory category (things to know before purchase):
[0101] 1. Specifications: What is the biological sleep mode?
[0102] 2. Promotions: Free EMI, 20k promotion plan, HDFC Bank cashback, ICICI cashback
[0103] Simultaneous category:
[0104] 1. Physical store: Where to buy
[0105] 2. Manufacturer information: Cancellation, return and replacement policy, warranty policy
[0106] 3. Related products: AR18BY5APWK, AR24BY4YBWK
[0107] Post - category:
[0108] 1. Usage / Maintenance: AC filter cleaning, automatic AC cleaning, E121 error, how to use the good sleep mode?
[0109] Figure 16 Shows a comparison of classifying one or more hyperlinks in Document 303 between the prior art and the present disclosure according to an embodiment of the present disclosure.
[0110] Refer to Figure 16 , the user asks the intelligent assistant about the Statue of Liberty. In response, the intelligent assistant actively asks the user about unknown preparatory categories and uses embedded hyperlinks to provide answers (i.e., no need to search on the web again).
[0111] In this way, the present disclosure classifies hyperlinks in a more efficient manner. For example, the hyperlinks in document 303 are classified according to the user accessing document 303 and / or the content of document 303.
[0112] Although specific language has been used to describe the present disclosure, no limitation is intended thereby. It will be apparent to those skilled in the art that various operational modifications may be made to the method so as to implement the inventive concept taught herein.
[0113] The accompanying drawings and the foregoing description give examples of embodiments. Those skilled in the art will understand that one or more of the described elements may be well combined into a single functional element. Alternatively, a particular element may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, the order of the processes described herein may be changed and is not limited to the manner described herein.
[0114] In addition, the acts of any flowchart need not be implemented in the order shown; nor is it necessary to perform all the acts. Moreover, those acts that are not dependent on other acts may be performed in parallel with other acts. The scope of the embodiments is in no way limited by these specific examples. Many variations are possible, such as differences in structure, dimensions, and use of materials, whether or not explicitly given in the specification. The scope of the embodiments is at least as broad as the scope given by the appended claims.
[0115] Although the present disclosure has been shown and described with reference to various embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. A method for classifying one or more hyperlinks in a document, the method comprises: identifying the one or more hyperlinks 101 in the document based on an analysis of the text strings of the document; analyzing the surrounding text strings around each of the one or more hyperlinks 103; and classifying the one or more hyperlinks into at least one of a plurality of predefined categories based on the analysis of the surrounding text strings around each of the one or more hyperlinks 105.
2. The method according to claim 1, further comprises: extracting user information associated with a user accessing the document; and selecting the one or more hyperlinks among the classified one or more hyperlinks based on the user information.
3. The method according to claim 2, wherein the user information is extracted from a memory, wherein at least one of the user's user profile information or browsing history information is stored in the memory.
4. The method according to any one of claims 2 and 3, wherein the user information includes at least one of the following items: the user's profile information, the user's demographic information, the user's education level, published documents written by the user, and documents uploaded by the user.
5. The method according to any one of claims 1 to 4, wherein the plurality of predefined categories include at least two of a preparatory category, a simultaneous category, or a post category.
6. The method according to any one of claims 1 to 5, wherein classifying the one or more hyperlinks further comprises: classifying the one or more hyperlinks using the surrounding text strings without referring to the one or more hyperlinks.
7. The method according to any one of claims 1 to 6, further comprises: generating a list of link representations corresponding to the classified one or more hyperlinks based on the surrounding text strings without using the content of the one or more hyperlinks; and displaying the list of link representations on a graphical user interface GUI.
8. The method according to claim 7, wherein the operation of displaying the list of link representations further comprises: creating a taxonomy, wherein the taxonomy includes a list of concepts related to the text strings of the document; arranging the classified one or more hyperlinks in a predefined order based on the taxonomy; and displaying the list of link representations based on the arrangement of the classified one or more hyperlinks in the predefined order.
9. The method according to any one of claims 1 to 8, wherein the document corresponds to one of a text document, a web document, an application-based document, or a cloud-based document.
10. A system for classifying one or more hyperlinks in a document, the system comprises: an identification unit 210 configured to identify the one or more hyperlinks in the document based on an analysis of the text strings of the document; An analysis unit 212, configured to analyze the surrounding text strings around each of the one or more hyperlinks; and A classification unit 214, configured to classify the one or more hyperlinks into at least one of a plurality of predetermined categories based on the analysis of the surrounding text strings around each of the one or more hyperlinks.
11. The system according to claim 10, further comprising: An extraction unit 216, configured to extract user information associated with a user accessing the document; and A selection unit 218, configured to select the one or more hyperlinks from the classified one or more hyperlinks based on the user information.
12. The system according to claim 11, wherein the extraction unit 216 is further configured to extract the user information from a memory, where at least one of user profile information or browsing history information of the user is stored in the memory.
13. The system according to any one of claims 10 to 12, wherein for classifying the one or more hyperlinks, the classification unit 214 is further configured to: classify the one or more hyperlinks using the surrounding text strings without referring to the one or more hyperlinks.
14. The system according to any one of claims 10 to 13, further comprising: A generation unit 220, configured to generate a list of link representations corresponding to the classified one or more hyperlinks based on the surrounding text strings without using the content of the one or more hyperlinks; and A display unit 222, configured to display the list of link representations on a graphical user interface GUI.
15. The system according to claim 14, wherein for displaying the list of link representations, the generation unit 220 is further configured to: create a taxonomy, where the taxonomy includes a list of concepts related to the text strings of the document, and arrange the one or more classified hyperlinks in a predefined order based on the taxonomy, and wherein the display unit 222 is configured to display the list of link representations based on the arrangement of the one or more classified hyperlinks in the predefined order.