Text selection method and device, equipment, storage medium and program product
By obtaining interface coordinates in the text selection operation and determining the second selected text from the global text content based on semantic analysis, the problem of inaccurate text selection in the prior art is solved, and the accuracy and efficiency of text selection are improved.
Patent Information
- Application Number
- CN202410098550.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the text selection method relies on touch operations, resulting in differences between the selected text and the text expected to be selected by the user object, the operation accuracy and user experience are low, and the human-computer interaction efficiency is low.
By obtaining the interface coordinates of the text selection operation, the first selected text is determined from the global text content based on the semantic dimension, and using this as a reference to determine the second selected text with which it has a semantic inclusion relationship, and the semantic analysis process is used to intelligently adjust the text selection range.
Improve the accuracy and efficiency of text selection, avoid the problem of inaccurate text selection caused by touch operation deviation, and improve the human-computer interaction experience.
Smart Images

Figure CN120373256A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a method, apparatus, device, storage medium, and program product for text selection. Background Art
[0002] Text selection refers to the process of marking the content selected by the user from the given text based on the trigger of the user under the condition of the given text, so as to realize the differential presentation of part of the information. Text selection is usually applied to various text-related scenarios such as text viewing scenarios, information extraction scenarios, search engine scenarios, etc.
[0003] In related technologies, text selection is usually implemented based on touch operations. The user needs to select a single character by clicking on the screen or select multiple character contents by dragging to determine the selected text composed of characters.
[0004] The above method can only perform text selection according to the user's operation on the given text. When there are deviations in the operation or some character recognition errors occur, it is easy to make a difference between the selected text and the text expected to be selected by the user. Therefore, there are great limitations in terms of operation accuracy and user experience, and the human-computer interaction efficiency is low. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, device, storage medium, and program product for text selection, which can intercept a second selected text from the global text content through the semantic dimension under the condition of text selection operation, so as to intelligently adjust the text selection range with the help of the semantic analysis process and improve the text selection efficiency. The technical solutions are as follows.
[0006] On the one hand, a method for text selection is provided, and the method includes:
[0007] Obtain the interface coordinates of the text selection operation, where the interface coordinates are the interface position information determined on the text display interface based on the text selection operation, and partial text content in the global text content is displayed on the text display interface;
[0008] Determine a first selected text from the global text content based on the interface coordinates;
[0009] Taking the first selected text as a reference, determine a second selected text from the global text content that has a semantic inclusion relationship with the first selected text; the semantic inclusion relationship is used to represent that the second selected text includes the first selected text, and the second selected text is the text obtained from the semantic dimension based on the first selected text.
[0010] On the other hand, a text selection device is provided, and the device includes:
[0011] An acquisition module, configured to acquire the interface coordinates of a text selection operation, where the interface coordinates are the interface position information determined on a text display interface based on the text selection operation, and partial text content in the global text content is displayed on the text display interface;
[0012] A determination module, configured to determine a first selected text from the global text content based on the interface coordinates;
[0013] The determination module is further configured to, with the first selected text as a reference, determine a second selected text from the global text content that has a semantic inclusion relationship with the first selected text; the semantic inclusion relationship is used to represent that the second selected text includes the first selected text, and the second selected text is a text obtained from the semantic dimension based on the first selected text.
[0014] On the other hand, a computer device is provided, and the computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the text selection method as described in any one of the embodiments of the present application above.
[0015] On the other hand, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the text selection method as described in any one of the embodiments of the present application above.
[0016] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text selection method as described in any one of the above embodiments.
[0017] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0018] The interface coordinates are obtained based on a text selection operation on the text display interface, so as to determine the first selected text from the global text content according to the interface coordinates. Furthermore, with the first selected text as a reference, the second selected text is determined from the global text content according to the semantic dimension. The second selected text is the text obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the problem of the selection limitation of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of the text selection operation, thereby intelligently adjusting the text selection range through the semantic analysis process and improving the text selection efficiency. Brief Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of the implementation environment provided by an exemplary embodiment of the present application;
[0021] Figure 2 It is a flowchart of the text selection method provided by an exemplary embodiment of the present application;
[0022] Figure 3 It is a flowchart of the coordinate conversion provided by an exemplary embodiment of the present application;
[0023] Figure 4 It is a flowchart of the text selection method provided by another exemplary embodiment of the present application;
[0024] Figure 5 It is a schematic diagram of the conversion of interface coordinates to global coordinates provided by an exemplary embodiment of the present application;
[0025] Figure 6 It is a flowchart of the text selection method provided by yet another exemplary embodiment of the present application;
[0026] Figure 7 It is a flowchart of the text selection method provided by still another exemplary embodiment of the present application;
[0027] Figure 8 It is a schematic diagram of the document object model tree, layout tree, and rendering tree provided by an exemplary embodiment of the present application;
[0028] Figure 9 It is a schematic diagram of the system framework of the text selection method provided by an exemplary embodiment of the present application;
[0029] Figure 10 It is a flowchart for extracting the context extraction part provided by an exemplary embodiment of the present application;
[0030] Figure 11 It is a processing timing diagram of the text selection method provided by an exemplary embodiment of the present application;
[0031] Figure 12 It is a flowchart for intelligently selecting text through the text selection system provided by an exemplary embodiment of the present application;
[0032] Figure 13 It is a structural block diagram of the text selection device provided by an exemplary embodiment of the present application;
[0033] Figure 14 It is a structural block diagram of the server provided by an exemplary embodiment of the present application. Detailed implementation manners
[0034] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0035] First, a brief introduction is given to the nouns involved in the embodiments of the present application.
[0036] Artificial Intelligence (AI): It is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called the large model or the basic model, and can be widely applied to the downstream tasks of major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0037] Machine Learning (ML): It is an interdisciplinary subject involving multiple fields, such as probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specializes in studying how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0038] In related technologies, text selection is usually based on touch operations. The user needs to select a single character by clicking on the screen or select multiple character contents by dragging to determine the selected text composed of characters. The above methods can only select text according to the user's operations on the given text. When there are deviations in the operations or some characters are misrecognized, it is easy to make a difference between the selected text and the text expected to be selected by the user. Therefore, there are large limitations in terms of operation accuracy and user experience, and the human-computer interaction efficiency is low.
[0039] In the embodiments of the present application, a text selection method is introduced. It can intercept the second selected text from the global text content through the semantic dimension under the condition of text selection operations, so as to intelligently adjust the text selection range through the semantic analysis process and improve the text selection efficiency. This text selection method can be applied to various text-related scenarios such as intelligent question-and-answer systems, event analysis scenarios, and document operation scenarios. The embodiments of the present application do not limit this.
[0040] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the interface coordinates, global text content, etc. involved in the present application are obtained under full authorization.
[0041] Secondly, the implementation environment involved in the embodiments of the present application is described. The text selection method provided by the embodiments of the present application can be implemented independently by the terminal, or by the server, or by data interaction between the terminal and the server. The embodiments of the present application do not limit this. Optionally, taking the interaction between the terminal and the server to execute the text selection method as an example for description.
[0042] Schematically, please refer to Figure 1, in this implementation environment, the terminal 110 and the server 120 are involved, and the terminal 110 and the server 120 are connected through the communication network 130.
[0043] In some embodiments, the terminal 110 has a text acquisition function and a text display function. Through the text acquisition function, text content can be acquired, and through the text display function, the acquired text content can be displayed so that the user can view, edit, and other operations on the text content.
[0044] Optionally, the terminal 110 implements the text display function through a text display interface, that is, the text content is displayed through the text display interface.
[0045] Schematically, taking the acquisition and display of global text content as an example, the text display interface is used to display the global text content, and the global text content represents all the text content that needs to be displayed through the terminal 110; due to the display limitations of the terminal 110, currently only partial text content in the global text content is displayed on the text display interface, that is, part of the text content is displayed through the text display interface of the terminal 110.
[0046] Schematically, the terminal 110 can receive a text selection operation through the text display interface. The text selection operation is used to represent the text selection process for the partial text content displayed in the text display interface. For example, long-pressing at least one text in the partial text content to implement the text selection operation, or double-clicking at least one text in the partial text content to implement the text selection operation, etc.
[0047] Optionally, the terminal 110 determines the interface coordinates based on the text selection operation received on the text display interface. The interface coordinates are the coordinates determined relative to the screen displayed by the terminal 110, and are the interface position information determined based on the text selection operation on the text display interface, which can represent the position of the text selection operation relative to the screen.
[0048] In an alternative embodiment, the terminal 110 sends the interface coordinates to the server 120 through the communication network 130; the server 120 will determine the first selected text from the global text content based on the interface coordinates.
[0049] Among them, the first selected text is the text content determined based on the text selection operation. The interface coordinates help to locate the position of the text selection operation relative to the global text content to determine the first selected text targeted by the text selection operation.
[0050] In an alternative embodiment, taking the first selected text as a reference, the second selected text having a semantic inclusion relationship with the first selected text is determined from the global text content.
[0051] Among them, the semantic inclusion relationship is used to represent that the second selected text contains the first selected text, and the second selected text is obtained from the semantic dimension based on the first selected text.
[0052] Illustratively, according to the semantic relationship of the first selected text with respect to the global text content, the second selected text is determined from the global text content. The second selected text includes at least the first selected text, and may also include some text before the first selected text, or may also include some characters after the first selected text, etc.
[0053] It should be noted that the above terminals include but are not limited to mobile terminals such as mobile phones, tablet computers, portable laptop computers, intelligent voice interaction devices, intelligent household appliances, vehicle-mounted terminals, etc., and can also be implemented as desktop computers, etc.; the above servers can be independent physical servers, or can be a server cluster or distributed system composed of multiple physical servers, and can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0054] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, application programs, and networks within a wide area network or local area network to achieve data calculation, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, and can form a resource pool, which can be used on demand and is flexible and convenient.
[0055] In some embodiments, the above server can also be implemented as a node in a blockchain system.
[0056] Combined with the above noun introduction and application scenarios, the text selection method provided by this application is described. Taking this method applied to a server as an example, as Figure 2 shown, this method includes the following steps 210 to step 230.
[0057] Step 210, obtain interface coordinates.
[0058] Among them, the interface coordinates are interface position information determined on the text display interface based on a text selection operation.
[0059] Schematically, the text display interface is an interface displayed on a terminal for presenting text content. For example: the text display interface is a web page; or, the text display interface is a news viewing interface within a news application; or, the text display interface is a text editing interface within a text application; or, the text display interface is a novel viewing interface within a reading application, etc.
[0060] Optionally, when the terminal displays the text display interface, it can also receive a text selection operation, which is used to select the text content displayed on the text display interface; the text selection operation is implemented as at least one of multiple trigger operation types such as a long - press selection operation, a double - click selection operation, a drag - and - drop selection operation, etc.
[0061] Schematically, the text display interface is currently displaying the content of page 3 of a paper. A long - press operation on a certain position on the text display interface is regarded as a text selection operation.
[0062] Optionally, if the terminal receives a text selection operation on the text display interface, it will determine the interface coordinates based on the screen interface displayed by the terminal. The interface coordinates are the interface position information determined by integrating the text display interface and the text selection operation, and can characterize the position of the text selection operation relative to the text display interface.
[0063] Schematically, if a text selection operation is received at point A on the text display interface, the interface coordinates determined based on the text display interface and the text selection operation are used to describe the interface position information of point A relative to the text display interface, that is: the interface coordinates are the coordinates of point A on the text display interface.
[0064] Among them, the local text content in the global text content is displayed on the text display interface.
[0065] Schematically, as an interface for displaying text content, the text display interface can present the global text content, where the global text content is used to represent all the text content that can be displayed on the current interface; the local text content is used to represent a part of the global text content.
[0066] Optionally, the global text content consists of multiple characters; or, the global text content consists of multiple characters and at least one image, etc.; based on this, the local text content in the global text content may consist of at least one character, or may consist of at least one character and at least one image, etc., which is not limited here.
[0067] For example: The text display interface is a paper viewing interface, the global text content is the paper D currently being viewed, and the local text content is part of the text content in paper D. For example, the local text content is the content on page 3 of paper D; or, the local text content is the content on pages 3 and 4 of paper D; or, the local text content is part of the content on page 3 and part of the content on page 4 of paper D, etc.
[0068] Optionally, at any time, the text display interface may display the global text content, or may only display the local text content in the global text content. That is: the text display interface has display limitations and may only present part of the text content in the global text content.
[0069] In some embodiments, the terminal determines the interface coordinates based on the text display interface and the text selection operation, and sends the interface coordinates to the server. The server obtains the interface coordinates and performs the text selection process based on the interface coordinates.
[0070] By using the server to execute the text selection method, it is beneficial to reduce the data processing load of the terminal, avoid the cumbersome processes of the terminal simultaneously performing interface rendering, coordinate generation, and coordinate analysis. By having the server undertake the coordinate analysis process, the efficiency of the terminal for interface rendering is improved, and it is also beneficial to the accuracy of coordinate generation.
[0071] In some embodiments, the terminal determines the interface coordinates based on the text display interface and the text selection operation, and thus performs the text selection process based on the obtained interface coordinates.
[0072] By using the terminal itself to execute the text selection method, it is beneficial for the terminal to focus on its own data processing process, avoid the problem of slow coordinate analysis efficiency caused by interacting with the server, and perform the coordinate analysis process through some processing engines while performing interface rendering, which is beneficial to more efficiently determining the selected text content based on the interface coordinates.
[0073] Step 220, determine the first selected text from the global text content based on the interface coordinates.
[0074] Schematically, the interface coordinates represent the position information of the text selection operation relative to the local text content.
[0075] In some embodiments, determine the first selected text from the local text content based on the interface coordinates.
[0076] Since the local text content is the text content in the global text content, the process of determining the first selected text can also be regarded as a process of determining from the global text content.
[0077] Optionally, traverse the local text content and determine the region bounding box corresponding to each text region. The text region is used to represent the region occupied by the characters corresponding to the rendering nodes on the text display interface.
[0078] Among them, the text region can be implemented as a character region composed of at least one character; or, the text region is implemented as a vocabulary region composed of at least one vocabulary; or, the text region is implemented as a sentence region composed of at least one sentence, etc. That is: the local text content can be pre-divided into multiple text regions according to at least one of the character region division method, the vocabulary region division method, or the sentence region division method, and each text region corresponds to a region bounding box.
[0079] Optionally, compare the global coordinates with the region bounding boxes corresponding to the multiple text regions, determine the region bounding box where the global coordinates are located, and use the text within the region bounding box as the first selected text. The first selected text may be implemented as at least one character, at least one vocabulary, or at least one sentence.
[0080] Through the above process, the first selected text targeted by the text selection operation can be determined from the local text content more quickly according to the interface coordinates, improving the acquisition efficiency of the first selected text; considering that the text selection operation usually has a small region range, the first selected text is usually located on the text display interface, and obtaining the text content with the help of the interface coordinates can also ensure good accuracy.
[0081] In some embodiments, the interface coordinates are converted to obtain global coordinates, and then the first selected text is determined from the global text content according to the global coordinates.
[0082] Optionally, determine the adjustment parameter based on the display situation between the global text content and the local text content, so as to adjust the interface coordinates with the adjustment parameter to obtain the global coordinates.
[0083] Schematically, the local text content is part of the global text content. The adjustment parameter is determined according to the display ratio and the document position corresponding to the local text content, and then the interface coordinates are adjusted with the adjustment parameter to obtain the global coordinates.
[0084] Among them, the display ratio is used to characterize the display situation of local text content relative to the original ratio. For example, 100% represents the original ratio, 120% represents the enlarged ratio, 60% represents the reduced ratio, etc.; the document position is used to characterize the position of the local text content in the global text content. Combining the display ratio and the document position can more accurately know the relative position of the displayed local text content in the global text content. The above process introduces the conversion of interface coordinates to global coordinates to determine the content of the first selected text through the global coordinates. According to the display situation between the global text content and the local text content, the adjustment parameters for adjusting the interface coordinates can be roughly determined. The adjustment parameters can establish the corresponding relationship between the interface coordinates and the global coordinates between the global text content and the local text content, obtain the global coordinates convenient for analyzing the global text content, and improve the accuracy of determining the first selected text from the global text content.
[0085] Step 230: Using the first selected text as a reference, determine the second selected text having a semantic inclusion relationship with the first selected text from the global text content.
[0086] Schematically, after obtaining the first selected text, based on the first relative position of the first selected text in the local text content and the second relative position of the local text content in the global text content, map the first selected text to the global text content to determine the second selected text from the global text content.
[0087] Among them, the first relative position is used to describe the position of the first selected text relative to the local text content, and the second relative position is used to describe the position of the local text content relative to the global text content; combining the first relative position and the second relative position can roughly determine the position of the first selected text relative to the global text content, and then determine the second selected text from the global text content according to the first selected text.
[0088] Among them, the semantic inclusion relationship is used to characterize that the second selected text includes the first selected text, and the second selected text is the text obtained from the semantic dimension based on the first selected text.
[0089] Optionally, in the process of mapping the first selected text to the global text content, multiple character contents before and after the first selected text will be determined through the global text content, and then the second selected text will be obtained by comprehensively considering the first selected text and the multiple character contents from the semantic dimension.
[0090] Schematically, the second selected text includes the first selected text, and the second selected text is obtained by expanding the text content based on the first selected text. That is: the second selected text is also an excerpt of the global text content, and can complete the semantic complement on the basis of the first selected text to more accurately reflect the semantic information of the text content where the first selected text is located.
[0091] For example: the first selected text is "apple", and the multiple character contents before "apple" include "I like to eat". The second selected text is determined by comprehensively analyzing "apple" and the multiple character contents from the semantic dimension. The second selected text is "I like to eat apples". Thus, while the text selection operation indicates a small part of the content, the second selected text that more conforms to the expressed semantic information can be obtained through the semantic analysis process.
[0092] It should be noted that the above is only a schematic example, and the embodiments of the present application are not limited thereto.
[0093] In summary, the interface coordinates are obtained according to the text selection operation on the text display interface, and then the first selected text is determined from the global text content according to the interface coordinates. Furthermore, with the first selected text as a reference, the second selected text is determined from the global text content according to the semantic dimension. The second selected text is the text obtained from the semantic dimension based on the first selected text. Therefore, the problem of the selection limitation of the first selected text can be avoided, and the second selected text including the first selected text can be intercepted from the global text content through the semantic dimension under the condition of the text selection operation, so as to intelligently adjust the text selection range through the semantic analysis process and improve the text selection efficiency.
[0094] In an optional embodiment, when determining the first selected text based on the interface coordinates, first convert the interface coordinates into global coordinates, and then determine the first selected text from the global text content based on the global coordinates. Schematically, as Figure 3 shown, the above Figure 2 shown step 220 can also be implemented as the following steps 310 to step 320.
[0095] Step 310, perform coordinate conversion on the interface coordinates to obtain global coordinates.
[0096] Schematically, when the text display interface displays local text content, the interface coordinates determined based on the text selection operation and the text display interface represent the positional relationship between the text selection operation and the local text content, and cannot represent the positional relationship between the text selection operation and the global text content. Therefore, after determining the interface coordinates, coordinate conversion needs to be performed on the interface coordinates to obtain the global coordinates representing the positional information between the text selection operation and the global text content.
[0097] That is: The global coordinates are used to represent the position information of the text selection operation relative to the global text content.
[0098] Optionally, the interface coordinates are coordinates determined based on the screen coordinate system corresponding to the terminal screen. Since the terminal screen is used to display the text display interface, the screen coordinate system can also be regarded as a coordinate system determined based on the text display interface.
[0099] Schematically, taking the upper left corner of the text display interface as the origin, the direction extending to the right from the origin is the positive direction of the horizontal axis (x-axis), and the direction extending downward from the origin is the positive direction of the vertical axis (y-axis); in addition, it can be stipulated that each pixel is a unit, so as to establish the screen coordinate system.
[0100] It should be noted that the selection of the origin, the selection of the positive direction of the axis, and the selection of the unit when constructing the screen coordinate system here are only schematic examples, and the embodiments of the present application do not limit this.
[0101] Optionally, the global coordinates are coordinates determined based on the global coordinate system corresponding to the global text content.
[0102] Schematically, taking the first text character of the global text content as the origin, the direction extending to the right from the origin is the positive direction of the horizontal axis (x-axis), the direction extending downward from the origin is the positive direction of the vertical axis (y-axis), and one pixel is a unit, so as to establish the global coordinate system; or, the global text content is a document, taking the upper left corner of the first page of the document as the origin, the direction extending to the right from the origin is the positive direction of the horizontal axis (x-axis), the direction extending downward from the origin is the positive direction of the vertical axis (y-axis), and one pixel is a unit, so as to establish the global coordinate system, etc.
[0103] It should be noted that the selection of the origin, the selection of the positive direction of the axis, and the selection of the unit when constructing the global coordinate system here are only schematic examples, and the embodiments of the present application do not limit this.
[0104] The above content has respectively described the global coordinate system and the screen coordinate system. Compared with the screen coordinate system, the global coordinate system has stronger unity, overcomes the problem that the interface coordinates may be different due to the corresponding screen coordinate systems of different terminals, and more accurately measures the position of the text selection operation relative to the global text content through the coordinate conversion process, which is beneficial to a more accurate analysis process based on the global coordinates.
[0105] In some embodiments, the adjustment parameter is determined based on the display situation between the global text content and the local text content, so as to adjust the interface coordinates with the adjustment parameter to obtain the global coordinates.
[0106] Schematically, the partial text content is part of the global text content. The adjustment parameter is determined according to the display ratio corresponding to the partial text content and the document position, and then the interface coordinates are adjusted with the adjustment parameter to obtain the global coordinates.
[0107] Among them, the display ratio is used to characterize the display situation of the partial text content relative to the original ratio. For example, 100% represents the original ratio, 120% represents the enlarged ratio, 60% represents the reduced ratio, etc.; the document position is used to characterize the position situation of the partial text content in the global text content. Combining the display ratio and the document position can more accurately know the relative position situation of the displayed partial text content in the global text content.
[0108] Step 320, determine the first selected text from the global text content based on the global coordinates.
[0109] Schematically, the global coordinates are the content obtained after coordinate transformation of the interface coordinates related to the terminal screen itself. Therefore, after obtaining the global coordinates, the global text content can be searched according to the global coordinates to obtain the first selected text.
[0110] In some embodiments, the global text content corresponds to multiple text regions, and each text region corresponds to a region bounding box and region attributes. The region bounding box is used to define the boundary for effective text selection of the characters within the text region, and the region attributes are used to describe the attributes of the unique text region, such as: position attributes, size attributes, etc.
[0111] Among them, the text region can be implemented as a character region composed of at least one character; or, the text region is implemented as a vocabulary region composed of at least one vocabulary; or, the text region is implemented as a sentence region composed of at least one sentence, etc. That is: the global text content can be pre-divided into multiple text regions according to at least one of the character region division method, the vocabulary region division method, or the sentence region division method, and each text region corresponds to a region bounding box respectively.
[0112] Based on the limitation of the region attributes, each text region in the global text content corresponds to a unique region bounding box. Therefore, the region bounding boxes corresponding to the multiple text regions can be obtained.
[0113] Schematically, each text character (such as "he", "I", etc.) represents a text region, and then the text character corresponds to a region bounding box; or, at least two characters forming a word represent a text region, and then the vocabulary corresponds to a region bounding box. For example: "apple" composed of two characters represents a text region, "hami melon" composed of three characters represents a text region, etc.
[0114] Optionally, each text area corresponds to area attributes such as area size and area position. The area size is used to characterize the size of the text area, and the area position is used to characterize the coordinate range of the text area relative to the global text content.
[0115] For example: The area size corresponding to text area A is 1*1, and the area position is represented as (1, 1)-(1, 2)-(2, 2)-(2, 1), which means that the text bounding box corresponding to text area A is a 1*1 square bounding box.
[0116] Optionally, the global coordinates are matched with the area bounding boxes respectively corresponding to multiple text areas, so as to determine the area bounding box where the global coordinates are located, and the text characters in this area bounding box are used as the first selected text.
[0117] The above determines the first selected text through the positional relationship between the area bounding box corresponding to the text area and the global coordinates. The area bounding box has the unique feature based on area attributes such as area position and area size. Comparing the area bounding boxes respectively corresponding to multiple text areas with the global coordinates can more accurately know the area bounding box where the global coordinates are located, and then more accurately know the text content of the first selected text with the text characters in the text area corresponding to this area bounding box, improving the acquisition accuracy of the first selected text.
[0118] In summary, the second selected text is the text obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the selection limitation problem of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of text selection operation, so as to intelligently adjust the text selection range through the semantic analysis process and improve the text selection efficiency.
[0119] In the embodiments of the present application, the conversion of interface coordinates to global coordinates is introduced, and then the content of the first selected text is obtained according to the global coordinates. The interface coordinates have certain limitations relative to the global text content, which is not conducive to obtaining accurate first selected text from the global text content. Converting the interface coordinates to global coordinates can show the position of the text selection operation relative to the global text content through a global coordinate system with high consistency and standardization, improving the accuracy of determining the first selected text from the global text content according to the global coordinates.
[0120] In an optional embodiment, according to the screen parameters corresponding to the terminal screen for displaying the text display interface, the interface coordinates are converted, and then the global coordinates are obtained. Schematically, as Figure 4 shown, step 310 shown above can also be implemented as steps 410 to 420 as follows. Figure 3 shown, step 310 shown above can also be implemented as steps 410 to 420 as follows.
[0121] Step 410, obtain the screen zoom parameter and the screen scroll parameter.
[0122] Schematically, based on the interface coordinates, the position information of the text selection operation relative to the text display interface is characterized. Therefore, the interface coordinates have strong limitations and cannot characterize the position of the text selection operation relative to the global text content. Therefore, in order to more accurately locate the position of the text selection operation in the global text content, it is necessary to convert the interface coordinates into global coordinates.
[0123] In the process of converting the interface coordinates into global coordinates, it is necessary to comprehensively consider the display relationship between the local text content and the global text content in the text display interface to determine the adjustment parameter for adjusting the interface coordinates, so as to convert the interface coordinates into global coordinates through the adjustment parameter.
[0124] In some embodiments, the adjustment parameter is implemented as the screen zoom parameter and the screen scroll parameter.
[0125] Among them, the screen zoom parameter is used to characterize the zoom ratio of the local text content relative to the global text content.
[0126] Schematically, the screen zoom parameter usually characterizes the zoom ratio of the text display interface, which can be expressed as scale and can also be called the zoom factor, representing the zoom ratio of the screen. For example, the screen zoom parameter is 100%, 80%, 160%, etc.
[0127] Among them, the screen scroll parameter is used to characterize the visible ratio of the local text content relative to the global text content.
[0128] Schematically, scrolling usually refers to the process in which an object causes the text display interface and the elements within the text display interface to scroll in its parent container by scrolling the mouse wheel, sliding the touchpad, or other gestures. The screen scroll parameter involves influencing factors such as the scroll container, the scroll direction, and the scroll position. The parent container is the container used to carry the global text content.
[0129] Optionally, the influencing factors involved in the screen scroll parameter are introduced as follows.
[0130] (1) Scroll container: A certain area on the text display interface may be a scroll container, which can contain more content than is fully visible on the screen. This scroll container can be the entire document window or a fixed area in the text display interface page.
[0131] (2) Scroll direction: The scroll direction can be vertical (up or down) or horizontal (left or right), and the specific scroll direction depends on the user's scroll gesture or program control.
[0132] (3) Scroll position: The scroll position represents the currently visible area in the scroll container (parent container). For example, in vertical scrolling, the scroll position can be represented as the top position of the scroll bar; or, the scroll position can be represented as the bottom position of the scroll bar; or, the scroll position can be represented by the distance from the top position of the scroll bar or the top edge of the scroll container to the top of the visible area; or, the scroll position can be represented by the number of pages; or, the scroll position can be represented by a scroll ratio. For example, a scroll position of 45% represents scrolling to the upper-middle part of the global text content, etc.
[0133] Schematically, taking the scroll position as the screen scroll parameter as an example, the screen scroll parameter can be represented as scroll, and can also be called the scroll offset, which represents the scroll position of the local text content relative to the global text content, and can also be called the scroll position of the page container where the local text content is located relative to the parent container where the global text content is located.
[0134] Step 420: Obtain global coordinates based on the interface coordinates, the screen scroll parameter, and the screen zoom parameter.
[0135] Schematically, adjust the interface coordinates with the screen scroll parameter and the screen zoom parameter to obtain the global coordinates.
[0136] In an optional embodiment, obtain candidate positioning coordinates based on the sum of the interface coordinates and the screen scroll parameter.
[0137] Optionally, since the screen scroll parameter represents the scrolling situation of the local text content relative to the global text content, the interface coordinates and the screen scroll parameter can be added to roughly determine the position of the text selection operation in the local text content, that is, determine the candidate positioning coordinates.
[0138] Schematically, using p to represent the trigger point position for the text selection operation, the interface coordinates can be represented as p_local, then the candidate positioning coordinates can be represented as "p_local + scroll".
[0139] In an optional embodiment, adjust the candidate positioning coordinates with the screen zoom parameter to obtain the global coordinates.
[0140] Optionally, since the screen zoom parameter represents the zoom ratio of the screen, it can also be regarded as the zooming situation of the local text content within the text display interface; therefore, the candidate positioning coordinates can be adjusted with the screen zoom parameter to restore the coordinate representation of the candidate positioning coordinates relative to the text display interface.
[0141] Schematically, when adjusting the candidate positioning coordinates according to the screen scaling parameter, the quotient of the candidate positioning coordinates and the screen scaling parameter is determined by using the candidate positioning coordinates as the dividend and the screen scaling parameter as the divisor.
[0142] For example, when using scale to represent the screen scaling parameter and adjusting the candidate positioning coordinates p_local + scroll according to the screen scaling parameter scale, (p_local + scroll) / scale is obtained.
[0143] The above process describes the adjustment of the interface coordinates according to the screen scaling parameter and the screen scrolling parameter. Quantifying the adjustment parameters for adjusting the interface coordinates as the screen scaling parameter and the screen scrolling parameter can consider the local text content relative to the text display interface through the screen scaling parameter, and can also consider the local text content relative to the global text content through the screen scrolling parameter. Combining the two parameters can make a more targeted adjustment to the interface coordinates and improve the accuracy of the global coordinates.
[0144] In some embodiments, considering that there may still be a distance between the page container where the local text content is located and the parent container where the global text content is located, the margin representing the distance between the page container and the parent container is obtained.
[0145] Schematically, the margin between the page container and the parent container is represented as margin.
[0146] In some embodiments, considering that there may still be a distance between the element border of the object element and the object element itself, CSS can also be used to obtain the distance between the internal content of each object element (including the rendering element among multiple object elements) and the corresponding element border, so as to obtain the inner margin of each object element; based on the fact that the CSS content is involved in the construction process of the rendering tree, the inner margin can also be referred to as the distance from the page container to the inside of the rendering tree. In the calculation process, it can be understood as: the distance between the element border of the rendering element corresponding to the rendering tree node targeted by the interface coordinates and the rendering element itself.
[0147] Schematically, the inner margin is represented as padding. The inner margin focuses more on the blank inside the rendering element, while the margin focuses more on the blank between the rendering element and its adjacent rendering elements. Both of them are important CSS properties for controlling the layout and the spacing between page elements.
[0148] In some embodiments, the margin, inner margin, screen scrolling parameter, and screen scaling parameter are combined to make a more accurate adjustment to the interface coordinates, so as to obtain the global coordinates.
[0149] Schematically, as shown in Formula 1 below, it is the calculation formula for converting interface coordinates into global coordinates.
[0150] Formula 1:
[0151] p_global = (p_local + scroll) / scale - margin - padding
[0152] Among them, p_global represents the global coordinates, p_local represents the interface coordinates, scroll represents the screen scrolling parameter, scale represents the screen scaling parameter, margin represents the margin, and padding represents the inner padding.
[0153] In an optional embodiment, as Figure 5 shown, the content of converting interface coordinates into global coordinates is described.
[0154] Schematically, as shown in Region 510, which includes a schematic diagram of the text display interface. The corresponding screen coordinate system is constructed with the upper left corner as the origin, the rightward direction as the positive x-axis direction, and the downward direction as the positive y-axis direction; the screen scrolling parameter is expressed through the x-axis and y-axis, that is: scroll = (scroll x, scroll y).
[0155] The screen scrolling parameter can be presented by the difference between Region 510 and Region 520. The text content in Region 520 is determined based on the Render Page and the Render Service. Among them, it corresponds to the scaled page container coordinate system, and the size of this scaled page container coordinate system changes according to the screen scaling situation.
[0156] The screen scaling parameter can be presented by the difference between Region 520 and Region 530. The text content in Region 530 is determined based on the Layout Page, Focus, and Focus Manager. Among them, it corresponds to the page container coordinate system, which is used to adjust from the physical size representation to the pixel size representation.
[0157] The margin and the inner padding can be presented by the difference between Region 530 and Region 540. Among them, the margin and the inner padding are fully considered to convert the interface coordinates represented by the screen coordinate system into the global coordinates represented by the global coordinate system.
[0158] In the above content, on the condition that the screen scaling parameter and the screen scrolling parameter are considered as adjustment parameters, the margin and the padding are additionally considered as adjustment parameters, so as to pay attention to the blank inside the rendered element through the padding and the blank between the rendered element and its adjacent rendered elements through the margin, which is beneficial to further strengthen the accuracy when adjusting the interface coordinates and obtain global coordinates more suitable for analyzing the global text content.
[0159] It should be noted that the above is only a schematic example, and the embodiments of the present application are not limited thereto.
[0160] To sum up, the second selected text is obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the selection limitation problem of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of the text selection operation, so as to intelligently adjust the text selection range through the semantic analysis process and improve the text selection efficiency.
[0161] In the embodiments of the present application, the content of adjusting the interface coordinates through the screen scaling parameter and the screen scrolling parameter to obtain the global coordinates is introduced. The screen scaling parameter represents the display ratio of the current text display interface, and the screen scrolling parameter represents the position of the local text content in the text display interface relative to the global text content. Combining the screen scaling parameter and the screen scrolling parameter can more accurately adjust the interface coordinates, facilitate obtaining more accurate global coordinates, and further improve the accuracy of determining the first selected text based on the global coordinates subsequently.
[0162] In an optional embodiment, when determining the first selected text from the global text content based on the interface coordinates, first obtain the global coordinates according to the interface coordinates, and then jointly determine the first selected text according to the global coordinates and the render tree. Schematically, as Figure 6 shown, the above Figure 3 step 320 shown can also be implemented as the following steps 610 to step 630.
[0163] Step 610, obtain the render tree.
[0164] Schematically, the render tree is a tree structure used to describe the structure of a Hyper Text Markup Language (HTML) document or an eXtensible Markup Language (XML) document and the rendering information of each element.
[0165] Optionally, the rendering tree consists of multiple rendering nodes, each of which represents a rendering element on the text display interface (which can also be referred to as a visible element, that is, an element that can be seen by the using object on the text display interface), and contains the rendering information related to the rendering of the rendering element; the rendering states and properties of the rendering elements reflected by each rendering node, and the rendering tree obtained by integrating multiple rendering nodes facilitate the layout and drawing of the rendering engine, and also facilitate the overall management of multiple rendering elements on the text display interface.
[0166] Schematically, the rendering elements represented by the rendering nodes include at least one of the following types.
[0167] (1) Page elements, such as: rendering elements such as paragraphs, headings, images, etc.; (2) Text elements, such as: characters in paragraphs, etc.; (3) Style elements, such as: colors, fonts, borders, etc.; (4) Layout elements: such as: layout attributes such as size, position, etc.
[0168] In some embodiments, the rendering tree is constructed based on the Document Object Model (DOM), and the text object model can also be referred to as the text object tree. Schematically, by parsing an HTML document or an XML document, the document is converted into a DOM tree, and the DOM tree represents the hierarchical structure of the document, including elements, attributes, and text nodes, etc.
[0169] Schematically, different from the DOM tree, the rendering tree only contains rendering nodes related to rendering and ignores nodes that do not affect the visual layout. Usually when constructing the rendering tree, first construct the DOM tree, and then calculate the style information of the rendering elements in the DOM tree, taking into account the Cascading Style Sheets (CSS) style rules, including inheritance and cascading rules; then, use the calculated style information to start constructing the rendering tree, and the rendering tree only includes rendering nodes that affect the page layout and rendering; in addition, each rendering node in the rendering tree will be assigned layout information, and the layout information includes the position and size of the rendering node on the text display interface, usually determined by considering the box model, floating, positioning and other properties of the element corresponding to the rendering node; finally, the interface rendering process can be executed according to the rendering tree, such as drawing the rendering elements represented by the rendering nodes in the rendering tree as pixels according to the layout information corresponding to each rendering node, so as to be displayed on the text display interface.
[0170] Optionally, the rendering nodes in the rendering tree can also be referred to as Render Objects. Each Render Object corresponds to a rendering element on the text display interface and contains the rendering information of the rendering element, such as size, color, position, etc. Through the construction process of the rendering tree, it is convenient to optimize the page layout by viewing the structure of the rendering tree and the rendering information of each Render Object.
[0171] Quantifying multiple rendering elements rendered and displayed on the text display interface in the global text content through the rendering tree is beneficial for ignoring the analysis of object elements that cannot be displayed on the text display interface, improving the element analysis efficiency and element rendering performance, and also helps to reduce the memory occupancy of invisible object elements, achieving the efficiency of dynamic update and the accuracy of element analysis.
[0172] Step 620: Traverse multiple rendering nodes in the rendering tree based on the global coordinates, and determine the rendering tree nodes in the multiple rendering nodes that have a matching relationship with the global coordinates.
[0173] Illustratively, after obtaining the global coordinates, multiple rendering nodes in the rendering tree can be traversed to find the rendering nodes corresponding to the global coordinates. That is: Determine the rendering tree nodes in the multiple rendering nodes that have a matching relationship with the global coordinates.
[0174] Among them, the matching relationship is used to represent that the global coordinates fall within the rendering area corresponding to the rendering tree node.
[0175] Optionally, during the process of traversing multiple rendering nodes in the rendering tree based on the global coordinates, determine the text areas corresponding to the multiple rendering nodes respectively.
[0176] Illustratively, each rendering node corresponds to a text area respectively, and the text area is the area where the rendering element represented by the rendering node is located.
[0177] Optionally, determine the text area that intersects with the global coordinates from the multiple text areas as the rendering area.
[0178] Among them, the rendering area corresponds to the rendering tree node; the rendering area is the text area corresponding to the rendering tree node.
[0179] Illustratively, when traversing multiple rendering nodes in the rendering tree based on the global coordinates, determine the text areas corresponding to the multiple rendering nodes respectively to find the text area that intersects with the global coordinates. The rendering node corresponding to this text area is the rendering tree node that has a matching relationship with the global coordinates. That is: The rendering tree node is used to represent the rendering node that has the above-mentioned matching relationship with the global coordinates.
[0180] Schematically, determine the text area where the global coordinates are located as the rendering area; each text area corresponds to a text bounding box, and the global coordinates are located within the text bounding box corresponding to the rendering area. Therefore, there is an intersection relationship between the text bounding box corresponding to the rendering area and the global coordinates.
[0181] Step 630, based on the rendering tree node, determine the first selected text from the global text content.
[0182] Among them, the first selected text is a rendering element rendered within the rendering area.
[0183] Schematically, the rendering tree node is one of the multiple rendering nodes in the rendering tree. After determining the rendering tree node, determine the characters rendered into the rendering area corresponding to this rendering tree node, so as to obtain the first selected text, and the first selected text includes at least one character.
[0184] In some embodiments, considering that the rendering tree usually pays more attention to the presentation information of the page rather than the direct text information, therefore, after determining the rendering tree node, the first selected text can be determined in combination with the text object tree.
[0185] Optionally, construct a text object tree based on the global text content.
[0186] Among them, the text object tree includes multiple object nodes, and the multiple object nodes respectively represent an object element. The multiple object elements are combined to obtain a document representing the global text content. The rendering tree is constructed based on some rendering elements among the multiple object elements in the text object tree. Therefore, there is a one-to-one correspondence between the multiple rendering nodes in the rendering tree and some object nodes in the text object tree.
[0187] Schematically, after determining the rendering tree node, the object node corresponding to the rendering tree node can be determined from the text object tree as the text object node, and the object element represented by this text object node is the first selected text.
[0188] That is: Once the rendering tree node is determined, based on the fact that each rendering node in the rendering tree usually corresponds to an object node in the DOM tree, therefore, the corresponding DOM node (text object node) can be found in the DOM tree through the association between the rendering tree and the DOM tree; then if it is determined that the rendering element represented by the rendering node is an element containing characters, the character content can be obtained by accessing the corresponding DOM node, and the first selected text is obtained.
[0189] It should be noted that the above is only a schematic example, and the embodiments of the present application are not limited thereto.
[0190] In summary, the second selected text is obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the problem of the selection limitation of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of text selection operation, so as to intelligently adjust the text selection range through the semantic analysis process and improve the text selection efficiency.
[0191] In the embodiments of the present application, the content of matching the global coordinates with the rendering tree to obtain the first selected text is introduced. Considering that the rendering elements displayed in the text display interface may change based on operations such as dragging and sliding, the rendering tree corresponding to the global text content is matched with the global coordinates, avoiding the interface structure limitation of the text display interface and improving the freedom of text selection; in addition, the rendering elements that can be displayed in the text display interface are more clearly managed through the rendering tree, and the global coordinates are respectively matched with multiple rendering nodes, so that the rendering elements represented by the rendering tree nodes with a matching relationship are used as the first selected text, improving the flexibility and accuracy of determining the selected text from the global text content based on the text selection operation.
[0192] In an optional embodiment, after determining the first selected text corresponding to the text selection operation, based on the first selected text, candidate text content including the first selected text is obtained from the global text content, and then the second selected text is determined through semantic dimension analysis. Schematically, as Figure 7 shown, the above Figure 2 illustrated embodiment can also be implemented as steps 710 to 740 as follows.
[0193] Step 710, obtain the interface coordinates.
[0194] Among them, the interface coordinates are the interface position information determined on the text display interface based on the text selection operation, and partial text content in the global text content is displayed on the text display interface.
[0195] In an optional embodiment, based on the trigger point position between the text selection operation and the text display interface, the interface coordinates are obtained.
[0196] Schematically, the text selection operation is a trigger operation for the text display interface displayed on the terminal screen. Therefore, after the text selection operation, the trigger point position for triggering the terminal screen can be determined by integrating the text selection operation and the text display interface.
[0197] Optionally, the terminal generates interface coordinates corresponding to the trigger point position according to the interface position situation of the trigger point position relative to the terminal screen, so as to represent the interface position information of the trigger point position relative to the text display interface through the interface coordinates.
[0198] In an optional embodiment, a screen coordinate system is established based on a preset origin in the text display interface.
[0199] The screen coordinate system is a coordinate system used to represent the interface coordinates.
[0200] Schematically, the interface coordinates are the coordinates determined in the screen coordinate system, and the preset origin is used to represent the coordinate origin in the screen coordinate system; the position of the preset origin is preset.
[0201] Optionally, taking the upper left corner of the text display interface as the preset origin, the positive x-axis direction of the screen coordinate system as the rightward extension of the preset origin, and the positive y-axis direction of the screen coordinate system as the downward extension of the preset origin as an example, a screen coordinate system is established to determine the coordinate expression of the interface coordinates.
[0202] In an optional embodiment, the interface coordinates are obtained based on the position information of the trigger point in the screen coordinate system.
[0203] The position information includes the offset distance of the trigger point relative to the preset origin and the offset direction of the trigger point relative to the preset origin. That is: based on the deviation distance and deviation direction of the trigger point relative to the preset origin in the text display interface, the interface coordinates are obtained.
[0204] The deviation distance and offset are values determined by comprehensively considering the horizontal deviation distance and vertical deviation distance of the trigger point. Optionally, if the screen coordinate system uses one pixel as a unit, the unit of the interface coordinates is expressed in pixels; or if the screen coordinate system uses a preset distance (such as 1 cm) as a unit, the unit of the interface coordinates is expressed in the preset distance as a unit, etc.
[0205] By comprehensively considering the offset distance and offset direction between the trigger point and the preset origin, the interface coordinates of the trigger point can be obtained more accurately, which is convenient for presenting the operation situation of the text selection operation more intuitively according to the interface coordinates.
[0206] Step 720: Determine the first selected text from the global text content based on the interface coordinates.
[0207] The first selected text is the text content determined based on the text selection operation.
[0208] In an optional embodiment, the interface coordinates are converted to obtain global coordinates.
[0209] The global coordinates are used to represent the position information of the text selection operation relative to the global text content.
[0210] The content of step 720 has been described in the above step 220 or steps 310 to 320, and will not be elaborated here.
[0211] Step 730: Taking the first selected text as a reference, determine candidate text content including the first selected text from the global text content according to a preset selection rule.
[0212] Illustratively, after obtaining the first selected text, based on the first relative position of the first selected text in the local text content and the second relative position of the local text content in the global text content, map the first selected text to the global text content to determine the second selected text from the global text content.
[0213] In some embodiments, determine the global coordinates corresponding to the interface coordinates. Taking the first selected text and the global coordinates as references, determine candidate text content including the first selected text from the global text content according to a preset selection rule.
[0214] Illustratively, the global coordinates are used to represent the position information of the first selected text relative to the global text content. After determining the global coordinates, other text content adjacent to the first selected text can be determined from the global text content, and the other text content can form candidate text content with the first selected text.
[0215] Wherein, the preset selection rule is used to obtain at least one character having a character adjacent relationship with the first selected text; the candidate text content is the text in the global text content.
[0216] Illustratively, the character adjacent relationship is used to represent that at least one character is adjacent to the first selected text. For example: I like to eat apples. If the first selected text is "eat", then at least one character having a character adjacent relationship is implemented as "like" and / or "apples"; or, the character adjacent relationship is used to represent that at least one character has an adjacent relationship with the first selected text based on adjacent characters, and the adjacent characters are determined based on the character order in the global text content. For example: I like to eat apples. If the first selected text is "eat", then at least one character having a character adjacent relationship is implemented as "like" and / or "apples", and "happy" can also be used as at least one character after determining "like", etc.
[0217] That is: taking the first selected text as a reference, candidate text content can be obtained from the global text content, and the candidate text content is the text content jointly composed of the first selected text and other text content before and after the first selected text.
[0218] Optionally, if the preset selection rule is implemented as a statement selection rule, then the statement where the first selected text is located can be used as the candidate text content.
[0219] Alternatively, the preset selection rule is implemented as a word count selection rule. Based on the first selected text, the text content before the first selected text by a preset number of words (e.g., 7 characters), the first selected text, and the text content after the first selected text by a preset number of words are jointly used to form the candidate text content, etc.
[0220] In an optional embodiment, a Document Object Model (DOM) tree is obtained.
[0221] Among them, the Document Object Model tree is a tree structure constructed based on the global text content. The text object model tree includes multiple object nodes, and each object node represents an object element. Multiple object elements constitute the global text content, and the object elements include at least one of various forms such as characters, images, and links.
[0222] In an optional embodiment, the text object model tree is traversed based on the first selected text to obtain the candidate text content.
[0223] Schematically, centered on the first selected text, multiple characters located before and after the first selected text are determined, and the multiple characters and the first selected text are combined in the character arrangement order in the global text content to form the candidate text content in the global text content.
[0224] Optionally, according to the traversal of the text object model tree, the characters with a preset number before the first selected text are determined, and the characters with a preset number after the first selected text are determined, so as to form the candidate text content.
[0225] For example: the first selected text is "is", the preset number is 5, the characters with a preset number before the first selected text are determined to be "system input", and the characters with a preset number after the first selected text are determined to be "object click", so the candidate text content is determined to be "system input is object click".
[0226] Optionally, according to the traversal of the text object model tree, the nearest punctuation mark before the first selected text is determined, and the nearest punctuation mark after the first selected text is determined, and the characters between the two punctuation marks are combined in the character arrangement order to form the candidate text content.
[0227] For example: the first selected text is "is", the nearest punctuation mark before the first selected text is determined to be "period". The nearest punctuation mark after the first selected text is determined to be "comma", and the characters between the two punctuation marks are combined in the character arrangement order to form the candidate text content - "system input is object click screen position", etc.
[0228] It should be noted that the above are only schematic examples, and the embodiments of the present application are not limited thereto.
[0229] In an alternative embodiment, a layout tree is obtained.
[0230] The layout tree is a tree structure constructed based on the rendering tree, which is used to more precisely display the layout of the rendering elements represented by multiple rendering nodes in the rendering tree, and is responsible for representing the exact position and size of each visible element in the text display interface.
[0231] The rendering tree is a tree structure constructed based on multiple rendering elements in the global text content.
[0232] Schematically, a DOM tree is constructed based on the global text content, and then a rendering tree is constructed based on the DOM tree, and further a layout tree is constructed based on the rendering tree.
[0233] The layout tree includes multiple layout nodes, and the hierarchical relationship between the multiple layout nodes is used to describe the layout relationship between the multiple rendering elements. The multiple rendering elements include the first selected text, and the first selected text is the rendering element corresponding to the rendering tree node in the rendering tree.
[0234] That is: the rendering tree node corresponds to the first selected text. When the rendering tree node is the result of matching the global coordinates with the rendering tree, it means that the rendering tree node corresponds to the global coordinates; when the rendering tree node is the result of matching the interface coordinates with the rendering tree, it means that the rendering tree node corresponds to the interface coordinates.
[0235] Schematically, the layout node represents the renderable elements in the global text content; the multiple rendering nodes in the rendering tree represent the rendering elements that need to be displayed in the global text content. The range of renderable elements may be larger than that of rendering elements because the renderable elements may include hidden elements, etc.
[0236] Schematically, by means of the hierarchical relationship between multiple layout nodes in the layout tree, the layout relationship between multiple renderable elements can be described, and the styles corresponding to multiple renderable elements can also be described.
[0237] In an alternative embodiment, the rendering tree node corresponding to the first selected text is matched with the layout tree, and the layout tree node corresponding to the rendering tree node is obtained from the multiple layout nodes.
[0238] Optionally, the rendering tree node is the rendering node obtained by matching the global coordinates with the rendering tree; or, the rendering tree node is the rendering node obtained by matching the interface coordinates with the rendering tree.
[0239] Schematically, after obtaining the global coordinates, the rendering tree nodes corresponding to the global coordinates are determined through the rendering tree; based on the one-to-one correspondence between the rendering nodes and some or all of the layout nodes, the rendering tree nodes can be matched with the layout tree, so as to obtain the layout tree nodes from multiple layout nodes, and the layout tree nodes are the layout nodes corresponding to the rendering tree nodes among the multiple layout nodes.
[0240] In an optional embodiment, under a preset selection rule, candidate text content including the first selected text is determined from the global text content through the layout tree nodes and the hierarchical relationship.
[0241] Schematically, the layout relationship within the layout tree is reflected through the hierarchical relationship between the layout nodes. Therefore, after determining the layout tree nodes, at least one layout node located above the layout tree node level and having a hierarchical relationship, and / or at least one layout node located below the layout tree node level and having a hierarchical relationship can be determined according to the hierarchical relationship within the layout tree. Thus, according to the layout elements represented by the determined at least one layout node, candidate text content including the first selected text is obtained.
[0242] Optionally, the number of selected layout nodes is preset; or, the number of selected layout nodes is randomly determined, etc., which is not limited here.
[0243] In an optional embodiment, as Figure 8 shown, it is a schematic diagram of the DOM tree 810, the layout tree 820, and the rendering tree 830.
[0244] Schematically, the global text content is usually implemented as an Office document, that is: the global text content is essentially content organized in a tree structure based on document data. Therefore, when constructing the DOM tree 810, the layout tree 820, and the rendering tree 830 based on the global text content, the top-level parent nodes all represent Document nodes, and their sub-nodes include body, section, column, etc. Each Dom node in each Dom tree will have corresponding layout nodes and rendering nodes, and what the user clicks is the rendering node in the rendering tree.
[0245] Optionally, first, the DOM tree 810 is constructed based on the global text content. For example, the root node of the DOM tree 810 is a file element, which is used to describe the global text content itself; in addition, the rendering tree 830 can be constructed based on the DOM tree 810, and the rendering elements represented by the rendering nodes in the rendering tree 830 are elements that can be rendered and displayed on the text display interface; furthermore, the layout tree 820 can be constructed based on the rendering tree 830, and the layout nodes in the layout tree 820 represent various layout relationships and layout information.
[0246] Step 740: Analyze the semantics of the candidate text content in the global text content, and determine the second selected text that has a semantic inclusion relationship with the first selected text from the global text content.
[0247] Illustratively, after obtaining the candidate text content, analyze the semantics of the candidate text content in the global text content. If the candidate text content has a clear semantic representation in the global text content, the candidate text content can be used as the second selected text; if the semantic representation of the candidate text content in the global text content is imperfect, character deletion can be performed on the candidate text content, or other characters in the global text content can be added before and after the candidate text content for supplementation, so as to obtain the second selected text, etc.
[0248] In an alternative embodiment, perform word segmentation on the candidate text content to obtain multiple text words.
[0249] Wherein, each text word includes at least one character; the multiple text words together form the candidate text content.
[0250] In an alternative embodiment, perform text recall on each of the multiple text words through a text retrieval model to obtain inverted lists corresponding to the multiple text words respectively.
[0251] Wherein, the text retrieval model is a model pre-trained through multiple sample documents, and each inverted list is used to represent a document set including the text word.
[0252] Optionally, when performing text recall on the multiple text words, input the multiple text words into the text retrieval model respectively, and the text retrieval model is a model pre-trained through multiple sample documents; determine the inclusion relationship between each text word and the multiple sample documents through the text retrieval model, so as to obtain inverted lists corresponding to the multiple text words respectively.
[0253] For example: The multiple sample documents include Sample Document 1, Sample Document 2, and Sample Document 3. Input the text word A into the text retrieval model pre-trained through the multiple sample documents, and determine that Sample Document 1 and Sample Document 2 include the text word A, and Sample Document 3 does not include the text word A. Then the inverted list corresponding to the text word A represents the document set composed of "Sample Document 1 and Sample Document 2".
[0254] In an alternative embodiment, determine at least one sample document including the multiple text words by integrating the multiple inverted lists.
[0255] Schematically, the inverted index corresponding to the text term A represents the document set composed of "sample document 1 and sample document 2", and the inverted index corresponding to the sample term B represents the document set composed of "sample document 1 and sample document 3". Then, the sample document 1 is a sample document that includes multiple text terms at the same time.
[0256] In an optional embodiment, at least one sample document is matched with the rendering tree corresponding to the global text content, and the second selected text is determined from the global text content.
[0257] Wherein, the rendering tree is a tree structure constructed based on multiple rendering elements in the global text content, and the rendering elements are the elements rendered on the text display interface.
[0258] Schematically, based on the fact that at least one sample document is a document that has been used to train the text retrieval model, the text terms in the sample document have strong relevance. Therefore, matching at least one sample document with the rendering tree corresponding to the global text content can, while ensuring that multiple text terms are mapped to the rendering tree, divide the selection of the global text content so as to determine the second selected text from the global text content.
[0259] Optionally, by marking the second selected text in ways such as highlighting and underlining, it can be visually presented on the text display interface, thereby improving the usage experience of the user.
[0260] That is: by matching to the rendering tree to form a selection area (the second selected text), it is mainly to flexibly determine the selection result on the text display interface, organically combine information retrieval and visualization, and facilitate easier understanding and interaction.
[0261] In an optional embodiment, the semantics of the candidate text content relative to the global text content is analyzed to obtain a semantic analysis result corresponding to the candidate text content.
[0262] Optionally, a semantic analysis model is obtained. The semantic analysis model is a neural network model pre-trained to have a semantic recognition function for sentences and / or document content, and is used to understand and interpret the semantic meaning of the text content to identify key information, emotions, context, etc. in the text content.
[0263] Schematically, the semantic analysis model is implemented as a Bidirectional Encoder Representations from Transformers (BERT), a Robustly optimized BERT approach, a Generalized Autoregressive Pretraining for Language Understanding (XLNet), etc.
[0264] Schematically, after obtaining the candidate text content, the candidate text content is input into the semantic analysis model, and the semantic analysis model will analyze the semantic dimension based on the candidate text content and output the semantic analysis result.
[0265] In some embodiments, the semantic analysis result reflects the semantic information of the candidate text content relative to the global text content.
[0266] Optionally, extract the text near the location of the candidate text content to form a context window, which can be several sentences, paragraphs or even the global text content before and after, and input the context window and the candidate text content into the semantic analysis model together to obtain a wider context and a more accurate semantic analysis result.
[0267] Optionally, extract the global feature representation corresponding to the global text content, and input the candidate text content and the global feature representation into the semantic analysis model together to obtain the semantic analysis result.
[0268] Optionally, when the semantic analysis model uses the attention mechanism, the attention weights of the semantic analysis model can be adjusted to make it pay more attention to the context of the location of the candidate text content, so that the semantic analysis model can pay more attention to the relevant parts during semantic analysis and output the semantic analysis result, etc.
[0269] In an optional embodiment, in response to the semantic analysis result meeting the preset conditions, the candidate text content is used as the second selected text.
[0270] Schematically, the preset condition is a condition that is preset and used to judge the semantic coherence of the semantic analysis result.
[0271] Optionally, the semantic analysis result is implemented as a semantic score, and the preset condition is implemented as a preset score threshold; the semantic score is used to characterize the semantic coherence score of the candidate text content in the global text content. When the semantic score reaches the preset score threshold, it is considered that the semantic analysis result meets the preset conditions, and the candidate text content coordinates the second selected text.
[0272] In an optional embodiment, in response to the semantic analysis result not meeting the preset condition, the candidate text content is adjusted based on the global text content to obtain the second selected text.
[0273] Optionally, the semantic analysis result is implemented as a semantic score, and the preset condition is implemented as a preset score threshold; when the semantic score does not reach the preset score threshold, it is considered that the semantic analysis result does not meet the preset condition, and at this time, the candidate text content needs to be adjusted to obtain the second selected text.
[0274] Schematically, character deletion is performed on the candidate text content according to the character arrangement order in the global text content, and the candidate text content after deletion is input into the semantic analysis model for semantic analysis until the second selected text is obtained after meeting the preset condition.
[0275] Schematically, characters adjacent to the candidate text content are determined from the global text content, and a character addition process is performed on the candidate text content according to the adjacent relationship between the characters and the candidate text content; then the candidate text content after addition is input into the semantic analysis model for semantic analysis until the second selected text is obtained after meeting the preset condition, etc.
[0276] It should be noted that the above are only schematic examples, and the embodiments of the present application are not limited thereto.
[0277] In summary, the second selected text is the text obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the selection limitation problem of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of text selection operation, thereby intelligently adjusting the text selection range through the semantic analysis process and improving the text selection efficiency.
[0278] In the embodiments of the present application, the content of analyzing from the semantic dimension based on the first selected text and obtaining the second selected text from the global text content is introduced. Taking the first selected text as a reference, candidate text content including the first selected text can be obtained through global coordinates, and then semantic analysis is performed based on the candidate text content. Under the premise of including the first selected text, with semantic coherence as the standard, the second selected text is obtained from the global text content. The second selected text is the text content analyzed from the semantic dimension based on the first selected text. It can, under the condition of selecting a single character or a small number of characters, expand and select other characters before and after the first selected text based on semantic factors to form the second selected text, making the second selected text more in line with the selection requirements of the user for text selection, improving the intelligence of text selection, and enhancing the text selection efficiency.
[0279] In an optional embodiment, the above text selection method is referred to as "an intelligent text selection method based on context word segmentation". As Figure 9 shown, it is a schematic diagram of the system framework of the text selection method, which includes two parts: a document engine 910 and a search middleware 920; the input of the text selection system is the interface coordinates. The document engine 910 analyzes based on the interface coordinates and obtains the context (the above-mentioned candidate text content), and then the search middleware 920 performs semantic dimension analysis based on the candidate text content to determine the second selected text.
[0280] Optionally, the document engine 910 and the search middleware 920 can be jointly deployed on the terminal, or jointly deployed on the server, or can be respectively deployed on the terminal and the server. For example, the document engine 910 is deployed on the terminal and the search middleware 920 is deployed on the server, etc. Schematically, the document engine 910 and the search middleware 920 are introduced as follows respectively.
[0281] (I) Document Engine 910
[0282] Schematically, the document engine 910 includes a coordinate system conversion part 911, a click test part 912, and a context extraction part 913. The coordinate system conversion part 911 is used to convert the interface coordinates into global coordinates; the click test part 912 is used to determine the first selected text; the context extraction part 913 is used to determine the candidate text content including the first selected text. Optionally, the coordinate system conversion part 911, the click test part 912, and the context extraction part 913 are respectively described as follows.
[0283] (1) Coordinate System Conversion Part 911
[0284] Schematically, according to the scaling of the screen (screen scaling parameter) and the scrolling of the screen (screen scrolling parameter), the interface coordinates are converted into global coordinates (device-independent coordinates), and the calculation formula is as shown in the above formula one.
[0285] Optionally, the coordinate system conversion part 911 inputs the global coordinates to the click test part 912.
[0286] (2) Click Test Part 912
[0287] Schematically, the click test part is used to receive the global coordinates and determine the first selected text based on the global coordinates.
[0288] Optionally, the gesture actions of the object on the screen are captured by a gesture recognizer, and the gesture actions include clicking, swiping, zooming, etc.; gesture recognition usually involves the process of registering the gesture recognizer to the RenderObject.
[0289] The registered gesture of the view is actually registered on the OffStageRenderObject. That is to say, a gesture recognizer is added to the view to capture gesture operations. The OffStageRenderObject is an object that is not displayed on the text display interface. For example, for the above object elements, in addition to the visible elements displayed on the view, there may also be some invisible elements, such as the elements blocked by the visible elements.
[0290] Optionally, if the registered gesture does not receive a callback, first consider whether this render object is hit; because only the hit render object can participate in gesture competition, and gesture competition is a process based on gesture conflict.
[0291] Schematically, when a gesture conflict occurs, first understand the hit test process of the render tree. The hit test process of the render tree is similar to that of the layout tree, both starting from the root node to the leaf node for hit testing. However, the details of the hit test of the render tree are relatively simple because there is no concept of hitting the interior. That is, a render object is either hit or not hit.
[0292] Schematically, when performing the hit test on the leaf node, it starts from the last leaf node added to the test. The later added leaf node means it is drawn later, and at the same position, it can block the drawing of the previous leaf node. That is to say, the drawing order of the render tree is from the bottom (root node) to the top (leaf node); while in the hit test, it goes from the leaf node to the root node, and the last added leaf node has the opportunity to be hit first in the hit test, so that the later added leaf node at the same position can block the drawing of the previous leaf node.
[0293] Optionally, generally the kIgnore method and the kOpaque method are used during the hit test process.
[0294] Among them, the kIgnore method means that during the hit test, the render object will be marked as ignored. This means that even if the position of the hit test is within the boundary of this render object, this object will not be considered a hit object and will not receive corresponding interaction events, such as click events or touch events. Thus, the hit test will continue to search for the next hittable object, and the current render object will not affect the result of the hit test.
[0295] The kOpaque method indicates that the rendering object is opaque, that is, during click testing, it will block the transmission of click events even if it itself is not clicked. If an opaque rendering object covers other sibling nodes, even if these sibling nodes are in the click area, they will still not be clicked due to the setting of the kOpaque method. Sibling nodes are other nodes that belong to the same level as the current rendering node and have the same parent node.
[0296] Optionally, it will continue to hit test its siblings only if its hit test behavior is kTranslucent.
[0297] In some embodiments, a hit test is performed on the rendering tree based on the global coordinates to determine the first selected text.
[0298] (3) Context Extraction 913
[0299] Illustratively, the first selected text obtained by performing a click test on the rendering tree based on the global coordinates often has only one word, so the result of the click test can be expanded to achieve the result that the user wants to select as much as possible.
[0300] Optionally, based on the first selected text of the click test, search forward and backward for directly adjacent text content. Only text content containing enough information can be the basis for accurate word segmentation. Considering speed and compliance, it can be limited to a preset length.
[0301] Illustratively, after the first selected text is determined, the first selected text is searched forward and backward with a preset length of 14 characters to determine the candidate text content.
[0302] like Figure 10 The figure shows the extraction flow chart of the context extraction part.
[0303] After the first selected text is determined based on the click test 1010, candidate text content is searched from the global text content based on the first selected text, and a real-time determination is made during the search whether the length limit is exceeded; when the length limit is not exceeded, the word-increasing search process is continued, and when the length limit is exceeded, a determination is made whether it is a text rendering object, that is, whether it is an object that can be rendered on the text display interface, such as: whether it is character content, etc.; when it is determined to be a character rendering object, the context can be expanded to obtain candidate text content including the first selected text.
[0304] Illustratively, the context information of the first selected text is continuously obtained in a deep traversal manner until the length exceeds the limit, and then the obtained context text content (candidate text content) and the node index of the rendering tree node in the rendering tree determined by the click test are input into the word segmentation module.
[0305] Among them, the node index is used to represent the rendering node corresponding to the rendered object clicked in the rendering tree. Through the node index, the first selected text can be quickly located.
[0306] (2) Search Middle Platform 920
[0307] Schematically, the search middle platform is used to execute the text segmentation (query, qu) process. Text segmentation is a basic method for text segmentation in the field of text retrieval, that is: an original document is cut into a word sequence according to the expectations of the user, and the word sequence includes multiple word segments. An inverted index chain is mapped through a single word segment, so as to perform text recall. The quality of the text segmentation result will directly affect the effect of text recall.
[0308] Schematically, the original document is "Introduction to Document Word Segmentation of Search Middle Platform", and multiple word segments include:
Search
Middle Platform
Document
Word Segmentation
Introduction
[0309] The above-mentioned cutting of the original document into 5 word segments corresponds to 5 inverted index chains (the document set where each word segment appears is the inverted index chain of that word segment). In the online recall part, each word segment recalls an inverted index chain, and the inverted index chains of multiple words can be merged to form the final recall result, that is, the second selected text is obtained.
[0310] In an optional embodiment, the display process of the text display interface is related to multiple engines, such as rendering engines, layout engines, editing engines, etc. Rendering engines, layout engines, and editing engines are key components related to fields such as Graphical User Interface (GUI) and web development, and usually play an important role in the presentation and interaction of user interfaces. The text display interface is a form of graphical user interface.
[0311] Schematically, the Rendering Engine is a component responsible for converting the content in a web page or application into a visual output to achieve the visual effect on the screen.
[0312] The Layout Engine is usually implemented as a part of the Rendering Engine and is responsible for processing and calculating the layout and position of each object element in the document to ensure that the visual presentation of the page meets the expected layout.
[0313] The Editing Engine is a component that performs text editing, text selection, etc. based on the operations of the user, and is involved in the interaction and processing of the user on the text box or text display interface. For example: the Editing Engine is responsible for capturing the text selection operation of the user and processing processes such as text selection, cursor movement, and image insertion.
[0314] In addition, the terminal implements the above content based on the terminal screen and can also perform processes with a large amount of data processing, such as text semantic analysis processes like word segmentation processing, through the background server. The terminal is implemented as a smart phone, a tablet computer, etc. The terminal screen is the screen equipped with the terminal (such as: a touch screen) and has various interface functions such as a text display function and a text selection function.
[0315] As Figure 11 shown, it is a processing sequence diagram of the text selection method, which involves the terminal screen 1110, the rendering engine 1120, the layout engine 1130, the editing engine 1140, and the background 1150. The background 1150 is used to store and process data and algorithms related to natural language processing (NPL).
[0316] Among them, the rendering engine 1120, the layout engine 1130, and the editing engine 1140 can be either components deployed on the terminal or components deployed on the server, and are not limited here.
[0317] First, the terminal screen 1110 captures an event based on the click process and sends it to the rendering engine 1120. The click process is regarded as the action form for implementing the text selection operation, and the captured event is regarded as capturing and determining the text selection operation; for example, the terminal screen 1110 captures the position where the user touches on the terminal screen through the touch operation recognition unit to determine the text range that the user wants to select.
[0318] The rendering engine 1120 determines the focus to obtain the interface coordinates indicated by the text selection operation.
[0319] After that, the layout engine 1130 determines the global coordinates based on the interface coordinates and determines the first selected text through a click test on the global coordinates, and sends the first selected text to the editing engine 1140.
[0320] Furthermore, the editing engine 1140 obtains the context based on the first selected text to obtain the candidate text content; then, the editing engine 1140 sends the candidate text content to the background 1150; for example: the editing engine 1140 performs data interaction with the background 1150 through the communication unit to send the candidate text content (text context information), and can also receive the word segmentation result through the communication unit.
[0321] The background 1150 performs word segmentation based on Natural Language Processing (NPL), thereby obtaining multiple text words as the word segmentation result; and sends the word segmentation result to the editing engine 1140; for example: the background 1150 receives the candidate text content through the NLP unit, performs word segmentation processing and returns the word segmentation result to the editing engine 1140;
[0322] The editing engine 1140 corrects the first selected text based on the word segmentation result, and then forms a selection area with the assistance of the layout engine 1130 and sends it to the rendering engine 1120. The selection area is used to represent the second selected text determined based on the word segmentation result; that is: the comprehensive editing engine 1140, layout engine 1130 and rendering engine 1120 determine the text content (the second selected text) that the user may want to select according to the word segmentation result and the touch position, and highlight the second selected text.
[0323] Finally, the rendering engine 1120 performs a rendering process such as drawing water droplets based on the second selected text, so as to display the second selected text with marks on the terminal screen 1110.
[0324] Optionally, operations such as copying, translating, and searching can also be performed on the intelligently selected second selected text, which is not limited here.
[0325] In an alternative embodiment, the above text selection method is applied to a text selection system. The text selection system can be deployed on a terminal, or in an application, or in a small program built into the application.
[0326] As Figure 12 shown, it is a flowchart of intelligently selecting text through the text selection system.
[0327] Step 1210, the user long-presses a character.
[0328] Schematically, the user long-presses a character to select part of the characters, that is, to determine the above first selected text.
[0329] Long-pressing to select characters is a basic ability in the document scenario. The document scenario is implemented as at least one of the word scenario, excel scenario, pdf scenario, and ppt scenario. That is: the text selection method proposed in the embodiments of the present application can be applied to multiple document scenarios to benefit many users, and the text selection method can be used in the above document scenarios to improve the efficiency of users in selecting character content. As shown in Table 1 below, it is the long-press count data statistically obtained for different document scenarios at different time periods.
[0330] Table 1
[0331] Date Word Excel Pdf PPT 20230605 1527.2W 5.6W 3366.4W 6.4W 20230604 1125.7W 3.4W 2777.4W 4.4W 20230603 921.6W 3.6W 2922.7W 3.6W 20230602 1146.6W 5.5W 3238.9W 4.8W 20230601 1254.1W 5.2W 3188.1W 4.9W
[0332] For example, at 05 June 2023, 15.272 million long - press operations were collected in a Word document, and at 05 June 2023, 0.056 million long - press operations were collected in an Excel document, etc.
[0333] Step 1220: Determine whether the function switch is turned on.
[0334] Schematically, the function switch is used to represent a text selection system for intelligent text selection; if the function switch is turned on, perform the following step 1230, that is, the text selection system can analyze based on the first selected text to perform the process of intelligent text selection; if the function switch is not turned on, directly determine that the first selected text is the text content selected by the user.
[0335] Step 1230: Return the complete sentence within this area.
[0336] Schematically, after determining the first selected text, use the complete sentence including the first selected text as the candidate text content for subsequent text analysis process.
[0337] Step 1240: Invoke the background word - segmentation function.
[0338] Schematically, after determining the candidate text content, analyze the candidate text content by invoking the background to perform word - segmentation processing on the candidate text content to obtain multiple word - segmentation results.
[0339] Step 1250: Update the selection area position according to the word - segmentation results by invoking the engine interface.
[0340] Schematically, perform semantic - dimension analysis through multiple word - segmentation results to determine a text selection area that better meets the user's needs based on the first selected text, and this text selection is the area corresponding to the second selected text.
[0341] It should be noted that the above is only a schematic example, and the embodiments of this application are not limited thereto.
[0342] In summary, the second selected text is the text obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the problem of the selection limitation of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of text selection operation, thereby intelligently adjusting the text selection range through the semantic analysis process and improving the text selection efficiency.
[0343] In the embodiments of the present application, considering the problems that may be encountered during the text selection operation such as long pressing by the user, with the help of the powerful word segmentation ability in the background, the second selected text that better meets the user's needs is automatically generated according to the context information of the first selected text. The second selected text is obtained by analyzing the first selected text from the semantic dimension. It not only belongs to the text in the global text content but also has strong semantic intelligence, which is conducive to significantly improving the accuracy and efficiency of the user's character selection.
[0344] In an optional embodiment, the text selection method is described below by taking the application scenario of document operation as an example.
[0345] Illustratively, the terminal is installed with a document application program, and a text display interface is displayed based on the triggering operation of the document application program. The text display interface is used to provide document operation functions such as document viewing, document editing, and text selection for the user.
[0346] Optionally, in response to receiving a text selection operation on the text display interface, the interface coordinates are determined. The interface coordinates are the interface position information determined on the text display interface based on the text selection operation, and partial text content in the global text content is displayed on the text display interface. For example, a paper A is displayed in the text display interface. The paper A has 16 pages, that is, the 16 - page paper is the global text content, and the partial text content displayed on the text display interface is the content of page 3. Then the text selection operation is a selection operation based on the text displayed on page 3.
[0347] The terminal can analyze the interface coordinates by itself or with the help of the server. Taking the server analysis as an example, the terminal sends the interface coordinates to the server. The server determines the first selected text from the global text content based on the interface coordinates, and then, taking the first selected text as a reference, determines the second selected text from the global text content that has a semantic inclusion relationship with the first selected text.
[0348] Among them, the semantic inclusion relationship is used to represent that the second selected text contains the first selected text, and the second selected text is the text obtained from the semantic dimension based on the first selected text.
[0349] Optionally, the second selected text is displayed in a highlighted form to prompt the user of the content they may expect to select; or, the second selected text is displayed in an underlined form to prompt the user of the content they may expect to select, etc.
[0350] By applying the text selection method to the document operation scenario, it is beneficial to provide a second selected text that better meets the expectations of the user during the document editing process, avoiding the limitation of solely relying on the user to perform text selection through text selection operations. By intercepting a second selected text with more semantic information from the global text content through the semantic analysis process based on the first selected text, it facilitates the user to perform processes such as copying and annotating the second selected text, thereby improving the efficiency of human-computer interaction.
[0351] In an optional embodiment, the text selection method applied to an intelligent question-and-answer system is used as an example for the following description.
[0352] Schematically, a terminal is deployed with an intelligent question-and-answer system (such as: an intelligent question-and-answer application program, an intelligent question-and-answer web engine, etc.). Based on the triggering operation of the intelligent question-and-answer system, a text display interface is displayed. The text display interface is used to provide a text question-and-answer function for the user. For example: the user asks a question to the system assistant, and the system assistant gives an answer text based on the question.
[0353] Optionally, in response to receiving a text selection operation on the text display interface, the interface coordinates are determined. The interface coordinates are the interface position information determined on the text display interface based on the text selection operation. The local text content in the global text content is displayed on the text display interface, and the local text content is implemented as part of the answer text. For example: multiple answer texts are obtained by answering multiple questions, and the local text content is one answer text; or, one answer text is obtained by answering multiple questions, and the local text content is part of the text in one answer text, etc.
[0354] The terminal can analyze based on the interface coordinates by itself, or it can also analyze the interface coordinates with the help of a server. Taking the terminal's self-analysis as an example, the terminal determines the first selected text from the global text content based on the interface coordinates, and then, taking the first selected text as a reference, determines the second selected text from the global text content that has a semantic inclusion relationship with the first selected text.
[0355] Among them, the semantic inclusion relationship is used to represent that the second selected text contains the first selected text, and the second selected text is the text obtained from the semantic dimension based on the first selected text.
[0356] Optionally, the second selected text is used as the next question text to ask the system assistant; or, the second selected text is used as a candidate question text for the user to select, so that the system assistant can be quickly asked questions, etc.
[0357] By applying the text selection method to the intelligent question-answering system, it is beneficial to provide a second selected text that better meets the expectations of the user during the intelligent question-answering process. By leveraging the user's selection process of the first selected text within the partial text content, the question text for the next intelligent conversation is analyzed, thereby facilitating the user to conduct intelligent conversations efficiently based on the second selected text and enhancing the human-computer interaction efficiency.
[0358] Figure 13 It is a structural block diagram of a text selection device provided by an exemplary embodiment of the present application. As Figure 13 shown, the device includes the following parts:
[0359] An acquisition module 1310, configured to acquire an interface coordinate, where the interface coordinate is interface position information determined based on a text selection operation on a text display interface, and the text display interface displays partial text content in the global text content;
[0360] A determination module 1320, configured to determine a first selected text from the global text content based on the interface coordinate;
[0361] The determination module 1320 is further configured to, with the first selected text as a reference, determine a second selected text having a semantic inclusion relationship with the first selected text from the global text content; the semantic inclusion relationship is used to represent that the second selected text includes the first selected text, and the second selected text is a text obtained from the semantic dimension based on the first selected text.
[0362] In an optional embodiment, the determination module 1320 is further configured to, with the first selected text as a reference, determine candidate text content including the first selected text from the global text content through a preset selection rule, where the preset selection rule is used to obtain at least one character having a character adjacency relationship with the first selected text; analyze the semantics of the candidate text content in the global text content, and determine the second selected text having the semantic inclusion relationship with the first selected text from the global text content.
[0363] In an alternative embodiment, the determining module 1320 is further configured to perform word segmentation on the candidate text content to obtain a plurality of text words; perform text recall on each of the plurality of text words through a text retrieval model to obtain an inverted posting list corresponding to each of the plurality of text words, where the text retrieval model is a model pre-trained through a plurality of sample documents, and the inverted posting list is used to represent a document set including the text word; comprehensively determine at least one sample document including the plurality of text words from the plurality of inverted posting lists; match the at least one sample document with a rendering tree corresponding to the global text content, and determine, from the global text content, the second selected text having the semantic inclusion relationship with the first selected text, where the rendering tree is a tree structure constructed based on a plurality of rendering elements in the global text content, and the rendering element is an element rendered on the text display interface.
[0364] In an alternative embodiment, the determining module 1320 is further configured to obtain a layout tree, where the layout tree is a tree structure constructed based on the rendering tree, the rendering tree is a tree structure constructed based on a plurality of rendering elements in the global text content, the layout tree includes a plurality of layout nodes, and the hierarchical relationship between the plurality of layout nodes is used to describe the layout relationship between the plurality of rendering elements, and the plurality of rendering elements include the first selected text, and the first selected text is a rendering element corresponding to a rendering tree node in the rendering tree; match the rendering tree node corresponding to the first selected text with the layout tree, and obtain a layout tree node corresponding to the rendering tree node from the plurality of layout nodes; under the preset selection rule, determine, from the global text content, the candidate text content including the first selected text through the layout tree node and the hierarchical relationship.
[0365] In an alternative embodiment, the determining module 1320 is further configured to perform coordinate conversion on the interface coordinates to obtain global coordinates, where the global coordinates are used to represent the position information of the text selection operation relative to the global text content; determine the first selected text from the global text content based on the global coordinates.
[0366] In an alternative embodiment, the determining module 1320 is further configured to obtain a rendering tree, which is a tree structure constructed based on a plurality of rendering elements in the global text content, and the rendering elements are elements rendered on the text display interface; the rendering tree includes a plurality of rendering nodes, and the plurality of rendering nodes correspond to the plurality of rendering elements one by one; traverse the plurality of rendering nodes in the rendering tree based on the global coordinates, and determine a rendering tree node in the plurality of rendering nodes that has a matching relationship with the global coordinates, where the matching relationship is used to represent that the global coordinates fall within the rendering area corresponding to the rendering tree node; based on the rendering tree node, determine the first selected text from the global text content, and the first selected text is a rendering element rendered within the rendering area.
[0367] In an alternative embodiment, the determining module 1320 is further configured to, during the process of traversing the plurality of rendering nodes in the rendering tree based on the global coordinates, determine text areas corresponding to the plurality of rendering nodes respectively, where the text areas are used to represent the areas occupied by the characters corresponding to the rendering nodes on the text display interface; determine a text area that intersects with the global coordinates from the plurality of text areas as the rendering area, and the rendering area corresponds to the rendering tree node.
[0368] In an alternative embodiment, the determining module 1320 is further configured to obtain a screen scaling parameter and a screen scrolling parameter, where the screen scaling parameter is used to represent the scaling condition of the local text content within the text display interface, and the screen scrolling parameter is used to represent the visible proportion of the local text content relative to the global text content; based on the interface coordinates, the screen scrolling parameter, and the screen scaling parameter, obtain the global coordinates.
[0369] In an alternative embodiment, the determining module 1320 is further configured to obtain candidate positioning coordinates based on the sum of the interface coordinates and the screen scrolling parameter; adjust the candidate positioning coordinates with the screen scaling parameter to obtain the global coordinates.
[0370] In an alternative embodiment, the obtaining module 1310 is further configured to obtain the interface coordinates based on the trigger point position between the text selection operation and the text display interface, and the interface coordinates are the position coordinates of the trigger point relative to the text display interface.
[0371] In an optional embodiment, the obtaining module 1310 is further configured to establish a screen coordinate system based on a preset origin in the text display interface; and obtain the interface coordinates based on the position information of the trigger point in the screen coordinate system, where the position information includes the offset distance of the trigger point position relative to the preset origin and the offset direction of the trigger point position relative to the preset origin.
[0372] In an optional embodiment, the determining module 1320 is further configured to use the first selected text as a reference to determine candidate text content including the first selected text from the global text content through a preset selection rule, where the preset selection rule is used to obtain at least one character having a character adjacent relationship with the first selected text; analyze the semantics of the candidate text content relative to the global text content to obtain a semantic analysis result corresponding to the candidate text content, where the semantic analysis result is used to represent the semantic information of the candidate text content in the global text content; and in response to the semantic analysis result meeting a preset condition, use the candidate text content as the second selected text.
[0373] In an optional embodiment, the determining module 1320 is further configured to, in response to the semantic analysis result not meeting the preset condition, adjust the candidate text content based on the global text content to obtain the second selected text.
[0374] In an optional embodiment, the determining module 1320 is further configured to delete characters from the candidate text content in accordance with the character arrangement order in the global text content, and perform semantic analysis on the candidate text content after deletion until the preset condition is met to obtain the second selected text; or determine characters adjacent to the candidate text content from the global text content, and add characters to the candidate text content according to the adjacent relationship between the characters and the candidate text content; perform semantic analysis on the candidate text content after addition until the preset condition is met to obtain the second selected text.
[0375] In summary, the second selected text is obtained from the semantic dimension based on the first selected text. Therefore, it can avoid the problem of the selection limitation of the first selected text, and intercept the second selected text including the first selected text from the global text content through the semantic dimension under the condition of text selection operation, so as to intelligently adjust the text selection range through the semantic analysis process and improve the text selection efficiency.
[0376] It should be noted that: The text selection device provided in the above embodiments is only illustrated by dividing the above function modules. In practical applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the text selection device provided in the above embodiments and the embodiments of the text selection method belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0377] Figure 14 FIG. shows a schematic structural diagram of a server provided by an exemplary embodiment of the present application. The server 1400 includes a central processing unit (CPU) 1401, a system memory 1404 including a random access memory (RAM) 1402 and a read only memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the central processing unit 1401. The server 1400 also includes a mass storage device 1406 for storing an operating system 1413, application programs 1414, and other program modules 1415.
[0378] The mass storage device 1406 is connected to the central processing unit 1401 through a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1406 and its associated computer-readable medium provide non-volatile storage for the server 1400. That is to say, the mass storage device 1406 can include computer-readable media (not shown) such as a hard disk or a compact disc read only memory (CD-ROM) drive.
[0379] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The above system memory 1404 and mass storage device 1406 can be collectively referred to as memory.
[0380] According to various embodiments of the present application, the server 1400 can also run on a remote computer on the network through a network such as the Internet. That is, the server 1400 can be connected to the network 1412 through a network interface unit 1411 connected to the system bus 1405, or in other words, the network interface unit 1411 can also be used to connect to other types of networks or remote computer systems (not shown).
[0381] The above-mentioned memory further includes one or more programs, and the one or more programs are stored in the memory and configured to be executed by the CPU.
[0382] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text selection method provided in each of the above method embodiments.
[0383] An embodiment of the present application further provides a computer-readable storage medium, on which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the text selection method provided in each of the above method embodiments.
[0384] An embodiment of the present application further provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text selection method described in any one of the above embodiments.
[0385] The foregoing are only optional embodiments of the present application, and are not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A text selection method, characterized in that, The method includes: Obtaining the interface coordinates of the text selection operation, where the interface coordinates are the interface position information determined based on the text selection operation on the text display interface, and partial text content in the global text content is displayed on the text display interface; Determining a first selected text from the global text content based on the interface coordinates; Taking the first selected text as a reference, determining a second selected text from the global text content that has a semantic inclusion relationship with the first selected text; the semantic inclusion relationship is used to represent that the second selected text includes the first selected text, and the second selected text is the text obtained from the semantic dimension based on the first selected text.
2. The method according to claim 1, characterized in that, The step of taking the first selected text as a reference and determining a second selected text from the global text content that has a semantic inclusion relationship with the first selected text includes: Taking the first selected text as a reference, determining candidate text content including the first selected text from the global text content through a preset selection rule, where the preset selection rule is used to obtain at least one character having a character adjacent relationship with the first selected text; Analyzing the semantics of the candidate text content in the global text content, and determining the second selected text from the global text content that has the semantic inclusion relationship with the first selected text.
3. The method according to claim 2, wherein The step of analyzing the semantics of the candidate text content in the global text content and determining the second selected text from the global text content that has the semantic inclusion relationship with the first selected text includes: Performing word segmentation on the candidate text content to obtain a plurality of text words; Performing text recall on each of the plurality of text words through a text retrieval model to obtain an inverted list corresponding to each of the plurality of text words, where the text retrieval model is a model pre-trained through a plurality of sample documents, and the inverted list is used to represent a document set including the text word; Comprehensively determining at least one sample document including the plurality of text words based on the plurality of inverted lists; Matching the at least one sample document with a rendering tree corresponding to the global text content, and determining the second selected text from the global text content that has the semantic inclusion relationship with the first selected text, where the rendering tree is a tree structure constructed based on a plurality of rendering elements in the global text content, and the rendering element is an element rendered on the text display interface.
4. The method according to claim 2, characterized in that The step of taking the first selected text as a reference and determining candidate text content including the first selected text from the global text content through a preset selection rule includes: Obtaining a layout tree, where the layout tree is a tree structure constructed based on the rendering tree, the rendering tree is a tree structure constructed based on a plurality of rendering elements in the global text content, the layout tree includes a plurality of layout nodes, the hierarchical relationship between the plurality of layout nodes is used to describe the layout relationship between the plurality of rendering elements, the plurality of rendering elements include the first selected text, and the first selected text is the rendering element corresponding to a rendering tree node in the rendering tree; Match the rendering tree node corresponding to the first selected text with the layout tree, and obtain the layout tree node corresponding to the rendering tree node from the multiple layout nodes; Under the preset selection rule, determine the candidate text content including the first selected text from the global text content through the layout tree node and the hierarchical relationship.
5. The method according to any one of claims 1 to 4, characterized in that, The determining the first selected text from the global text content based on the interface coordinates includes: Perform coordinate conversion on the interface coordinates to obtain global coordinates, where the global coordinates are used to represent the position information of the text selection operation relative to the global text content; Determine the first selected text from the global text content based on the global coordinates.
6. The method according to claim 5, characterized in that, The determining the first selected text from the global text content based on the global coordinates includes: Obtain a rendering tree, where the rendering tree is a tree structure constructed based on multiple rendering elements in the global text content, and the rendering elements are elements rendered on the text display interface; the rendering tree includes multiple rendering nodes, and the multiple rendering nodes correspond to the multiple rendering elements one by one; Traverse the multiple rendering nodes in the rendering tree based on the global coordinates, and determine a rendering tree node in the multiple rendering nodes that has a matching relationship with the global coordinates, where the matching relationship is used to represent that the global coordinates fall within the rendering area corresponding to the rendering tree node; Based on the rendering tree node, determine the first selected text from the global text content, where the first selected text is a rendering element rendered within the rendering area.
7. The method according to claim 6, wherein The traversing the multiple rendering nodes in the rendering tree based on the global coordinates and determining a rendering tree node in the multiple rendering nodes that has a matching relationship with the global coordinates includes: During the process of traversing the multiple rendering nodes in the rendering tree based on the global coordinates, determine the text areas corresponding to the multiple rendering nodes respectively, where the text area is used to represent the area occupied by the characters corresponding to the rendering node on the text display interface; Determine the text area that intersects with the global coordinates from the multiple text areas as the rendering area, and the rendering area corresponds to the rendering tree node.
8. The method according to claim 5, wherein The performing coordinate conversion on the interface coordinates to obtain global coordinates includes: Obtain a screen scaling parameter and a screen scrolling parameter, where the screen scaling parameter is used to represent the scaling situation of the local text content within the text display interface, and the screen scrolling parameter is used to represent the visible ratio of the local text content relative to the global text content; Based on the interface coordinates, the screen scrolling parameter, and the screen scaling parameter, obtain the global coordinates.
9. The method according to claim 8, wherein The obtaining the global coordinates based on the interface coordinates, the screen scrolling parameter, and the screen scaling parameter includes: Based on the sum of the interface coordinates and the screen scrolling parameter, obtain candidate positioning coordinates; Adjust the candidate positioning coordinates with the screen scaling parameter to obtain the global coordinates.
10. The method according to any one of claims 1 to 4, characterized in that The obtaining the interface coordinates includes: Based on the trigger point position between the text selection operation and the text display interface, obtain the interface coordinates, where the interface coordinates are the position coordinates of the trigger point relative to the text display interface.
11. The method according to claim 10, wherein The obtaining of the interface coordinates includes: Establish a screen coordinate system based on a preset origin in the text display interface; Based on the position information of the trigger point in the screen coordinate system, obtain the interface coordinates, where the position information includes the offset distance of the trigger point relative to the preset origin and the offset direction of the trigger point relative to the preset origin.
12. The method according to claim 1, wherein Taking the first selected text as a reference, determining a second selected text having a semantic inclusion relationship with the first selected text from the global text content includes: Taking the first selected text as a reference, determining candidate text content including the first selected text from the global text content through a preset selection rule, where the preset selection rule is used to obtain at least one character having a character adjacent relationship with the first selected text; Analyze the semantics of the candidate text content relative to the global text content to obtain a semantic analysis result corresponding to the candidate text content, where the semantic analysis result is used to characterize the semantic information of the candidate text content in the global text content; In response to the semantic analysis result meeting a preset condition, use the candidate text content as the second selected text.
13. The method according to claim 12, wherein The method further includes: In response to the semantic analysis result not meeting the preset condition, adjust the candidate text content based on the global text content to obtain the second selected text.
14. The method according to claim 13, characterized in that, The adjusting the candidate text content based on the global text content to obtain the second selected text includes: Delete characters from the candidate text content in accordance with the character arrangement order in the global text content, and perform semantic analysis on the candidate text content after deletion until the preset condition is met to obtain the second selected text; or, Determine characters adjacent to the candidate text content from the global text content, and add characters to the candidate text content according to the adjacent relationship between the characters and the candidate text content; perform semantic analysis on the candidate text content after addition until the preset condition is met to obtain the second selected text.
15. A text selection device, characterized in that, The device includes: An obtaining module, configured to obtain the interface coordinates of a text selection operation, where the interface coordinates are the interface position information determined based on the text selection operation on a text display interface, and a partial text content in the global text content is displayed on the text display interface; A determining module, configured to determine a first selected text from the global text content based on the interface coordinates; The determining module is further configured to, taking the first selected text as a reference, determine a second selected text having a semantic inclusion relationship with the first selected text from the global text content; the semantic inclusion relationship is used to characterize that the second selected text includes the first selected text, and the second selected text is a text obtained from a semantic dimension based on the first selected text.
16. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory. The at least one program is loaded and executed by the processor to implement the text selection method according to any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium. The at least one program is loaded and executed by a processor to implement the text selection method according to any one of claims 1 to 14.
18. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the text selection method according to any one of claims 1 to 14 is implemented.