Text selection method and extended reality system

By recognizing user gestures in virtual reality scenes and entering different text selection modes, the problems of high complexity and poor functional completeness of text selection operations in virtual environments are solved, and natural and efficient text selection is achieved.

CN120469612BActive Publication Date: 2025-09-12SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510955266.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing text selection interaction solutions in virtual environments have problems such as high operational complexity and poor functional completeness. Especially in the trend of natural interaction without devices, the limited gesture command set makes it difficult to meet multi-dimensional interaction needs, and the existing system lacks fault tolerance and correction mechanisms.

Method used

The text content is displayed in a virtual reality scene through a head-mounted display device, the user's gesture operation type is identified, and the corresponding text selection mode is entered, including single text segment manual selection mode, multiple text segment manual selection mode and text segment automatic selection mode, and the corresponding text selection rules are used to select the target text block.

Benefits of technology

It provides a fully functional text selection method, improves the naturalness and fluency of user interaction, reduces the user's cognitive burden, and meets multi-dimensional interaction needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469612B_ABST
    Figure CN120469612B_ABST
Patent Text Reader

Abstract

The present application relates to the field of gesture interaction technology, and in particular to a text selection method and an extended reality system, the method comprising: displaying text content on a virtual text page in a virtual reality scene through a head-mounted display device; identifying the type of gesture operation performed by the user in response to a first gesture operation performed by the user; entering a corresponding text selection mode according to the gesture operation type, the text selection mode including a single text segment manual selection mode, a multiple text segment manual selection mode, and a text segment automatic selection mode; in the corresponding text selection mode, selecting a target text block in the text content in response to a second gesture operation performed by the user. The present application can effectively solve the problems of high complexity and poor functional completeness of text interaction operations in existing virtual environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gesture interaction technology, and in particular to a text selection method and an extended reality system. Background Art

[0002] Current text selection interaction solutions in virtual environments face multiple technical challenges, primarily in the following areas: First, while physical hardware-based interaction methods provide clear command input, this approach not only reduces user immersion but also significantly increases the physical and mental burden on users, compromising the naturalness and fluidity of the interaction. Second, with the trend towards device-free natural interaction, the limited set of gesture commands struggles to meet multi-dimensional interaction needs. This presents a dilemma for system developers: either expand functionality by increasing interaction complexity, which increases the user's cognitive load; or simplify interaction functionality to maintain a basic experience, which in turn limits system functionality expansion. Furthermore, the inherent precision limitations of humans in spatial motion, coupled with the limitations of current head-mounted display hand tracking technology and the inadequacies of existing interaction systems in terms of fault tolerance and correction mechanisms, further exacerbate the difficulty of text selection tasks in virtual environments. Furthermore, existing solutions still have significant shortcomings in meeting the diverse needs of users through interoperability.

[0003] In summary, existing virtual text interaction solutions have significant limitations in terms of interaction complexity and functional completeness, and are unable to fully meet the actual needs of users. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a text selection method and an augmented reality system, which can effectively solve the problems of high complexity and poor functional completeness of text interaction operations in existing virtual environments.

[0005] In a first aspect, an embodiment of the present application provides a text selection method applicable to an extended reality system, the method comprising:

[0006] Displaying text content on a virtual text page in a virtual reality scene through a head-mounted display device;

[0007] In response to a first gesture operation performed by a user, identifying a type of gesture operation performed by the user;

[0008] Entering a corresponding text selection mode according to the gesture operation type, wherein the text selection modes include a single text segment manual selection mode, a multiple text segment manual selection mode, and an automatic text segment selection mode;

[0009] In the corresponding text selection mode, in response to the second gesture operation performed by the user, a target text block is selected in the text content using a corresponding text selection rule.

[0010] In a second aspect, an embodiment of the present application provides a text selection device applicable to an augmented reality system, the device comprising:

[0011] A display module, configured to display text content on a virtual text page in a virtual reality scene via a head-mounted display device;

[0012] an operation type identification module, configured to identify, in response to a first gesture operation performed by a user, a type of gesture operation performed by the user;

[0013] A mode entry module, configured to enter a corresponding text selection mode according to the gesture operation type; the text selection modes include a single text segment manual selection mode, a multiple text segment manual selection mode, and an automatic text segment selection mode;

[0014] The text selection module is configured to select a target text block in the text content using a corresponding text selection rule in response to a second gesture operation performed by the user in the corresponding text selection mode.

[0015] In a third aspect, an embodiment of the present application provides an extended reality system, comprising: a head-mounted display device, a camera device, and a processing center; the head-mounted display device is used to display text content on a virtual text page in a virtual reality scene;

[0016] The camera device is used to take a picture of the user's hand to obtain a gesture operation image;

[0017] The processing center is used to recognize the user's first gesture operation and / or second gesture operation based on the gesture operation image to implement a text selection method provided according to the first aspect of the present application.

[0018] The embodiments of the present application have the following beneficial effects:

[0019] The present application provides a text selection method suitable for an extended reality system, which displays text content on a virtual text page in a virtual reality scene via a head-mounted display device. In response to a first gesture operation performed by a user, the method identifies the type of gesture operation performed by the user. Based on the gesture operation type, the method enters a corresponding text selection mode. The text selection modes include a single text segment manual selection mode, a multiple text segment manual selection mode, and an automatic text segment selection mode. In the corresponding text selection mode, the method selects a target text block within the text content in response to a second gesture operation performed by the user. The present application selects text through different modes, providing excellent functional completeness. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A structural diagram of an extended reality system according to an embodiment of the present application is shown;

[0022] Figure 2 A flow chart of a text selection method according to an embodiment of the present application is shown;

[0023] Figure 3 A text input scenario diagram in a text selection method according to an embodiment of the present application is shown;

[0024] Figure 4 A schematic diagram of a coordinate system of a user operation space in a text selection method according to an embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram of a text selection step in a single text segment selection mode in a text selection method according to an embodiment of the present application is shown;

[0026] Figure 6 A schematic diagram of text input steps based on a virtual keyboard plane in a text selection method according to an embodiment of the present application is shown;

[0027] Figure 7 A schematic diagram of another text selection step in a single text segment selection mode in the text selection method according to an embodiment of the present application is shown;

[0028] Figure 8 A schematic diagram showing the steps of anchor point locking and adjusting in the text selection method according to an embodiment of the present application is shown;

[0029] Figure 9 A schematic diagram of a text selection step in a multi-text segment manual selection mode in a text selection method according to an embodiment of the present application is shown;

[0030] Figure 10 A schematic diagram of mechanical semantic segmentation levels in the automatic text segment selection mode in the text selection method according to an embodiment of the present application is shown;

[0031] Figure 11 A schematic diagram of a text selection step in a multi-text segment automatic selection mode in a text selection method according to an embodiment of the present application is shown;

[0032] Figure 12aA schematic diagram showing the palm in forehand and backhand determination in the text selection method according to an embodiment of the present application is shown;

[0033] Figure 12b A schematic diagram showing the angle of the palm in forehand and backhand determination in the text selection method according to an embodiment of the present application is shown;

[0034] Figure 12c The following diagram shows the definitions related to forehand and backhand determination and the anchor cursor diagram in the text selection method according to an embodiment of the present application;

[0035] Figure 13 A schematic diagram showing the steps of manually selecting front and back anchor points in the backhand fine-tuning mode in the text selection method according to an embodiment of the present application is shown;

[0036] Figure 14 A schematic diagram of selecting front and back eye movement anchor points in a backhand fine-tuning mode in a text selection method according to an embodiment of the present application is shown;

[0037] Figure 15 A schematic diagram showing the steps of decomposing and selecting a single text block to be selected in the text selection method according to an embodiment of the present application is shown;

[0038] Figure 16 A structural schematic diagram of a text selection device according to an embodiment of the present application is shown.

[0039] Description of main component symbols:

[0040] 110 - processing center; 120 - head-mounted display device; 130 - camera device; 210 - display module; 220 - operation type recognition module; 230 - mode entry module; 240 - text selection module. DETAILED DESCRIPTION

[0041] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0042] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0043] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.

[0044] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.

[0045] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0046] In order to effectively solve the problems of high complexity and poor functional completeness of text interaction operations in existing virtual environments, the present application provides a text selection method and an extended reality system.

[0047] This application provides an extended reality system, exemplary, such as Figure 1 As shown, the system includes: a processing center 110, a head-mounted display device 120 and a camera device 130. The head-mounted display device is used to display a virtual text page.

[0048] The head mounted display device is used to display text content on a virtual text page in a virtual reality scene. The head mounted display device 120 includes but is not limited to a head mounted display.

[0049] The camera device 130 is used to take a picture of the user's hand to obtain a gesture operation image. The camera device 130 includes but is not limited to a camera.

[0050] The processing center is used to recognize the user's first gesture operation and / or second gesture operation based on the gesture operation image to implement the text selection method of the embodiment of the present application.

[0051] For example, pinching and releasing your middle finger and thumb allows you to enter manual selection mode for a single text segment; pinching and moving the pinch point a certain distance along the Z axis allows you to enter manual selection mode for multiple text segments; and pinching and moving the pinch point a certain distance along the X, Y, or Z axis allows you to enter automatic text segment selection mode (semantic segmentation mode). The main idea of ​​this application is to select text using different gestures in different modes. Furthermore, fine-tuning gestures can be used to fine-tune the selected text.

[0052] The text selection method is described below with reference to some specific embodiments.

[0053] Figure 2 A flowchart of a text selection method according to an embodiment of the present application is shown. Exemplarily, the text selection method includes the following steps:

[0054] S100: Display text content on a virtual text page in a virtual reality scene through a head-mounted display device.

[0055] Exemplarily, text content is displayed on a virtual text page in a VR, AR, or MR scene displayed by a head-mounted display. For example, the virtual text page is a web page in the VR, AR, or MR scene.

[0056] S200 : In response to a first gesture operation performed by a user, identifying a type of gesture operation performed by the user.

[0057] S300 , entering a corresponding text selection mode according to the gesture operation type. The text selection modes include a single text segment manual selection mode, a multiple text segment manual selection mode, and an automatic text segment selection mode.

[0058] The embodiment of the present application includes multiple text selection modes. When the user performs different gesture operations, the system will enter different text selection modes after recognition.

[0059] Exemplarily, the text segment automatic selection mode includes a single text segment automatic selection mode and a multiple text segment automatic selection mode. According to the gesture operation type, the corresponding text selection mode is entered, including:

[0060] If the gesture operation type is the first gesture A combination type, the control system enters the single text segment manual selection mode; the gestures in the first gesture A combination type include continuously performing a pinch action, releasing the pinch action, or using the palm to pass through the virtual keyboard plane of the system;

[0061] If the gesture operation type is a first gesture B combination type, the control system enters a multi-text segment manual selection mode; the gesture in the first gesture B combination type includes performing a pinch action and dragging the pinch point a preset distance along the first coordinate axis toward the user's body;

[0062] If the gesture operation type is the first gesture C combination type or the first gesture D combination type, the control system enters the text segment automatic selection mode; the gesture in the first gesture C combination type includes performing a pinch action and dragging the pinch point a preset distance along the second coordinate axis or the third coordinate axis; the gesture in the first gesture D combination type includes performing a pinch action and dragging the pinch point a preset distance along the first coordinate axis.

[0063] S400 , in a corresponding text selection mode, in response to a second gesture operation performed by a user, selecting a target text block in the text content using a corresponding text selection rule.

[0064] In different text selection modes, different text selection rules are used to select the target text block.

[0065] In one embodiment, the text selection mode includes a single text segment manual selection mode, a multiple text segment manual selection mode, and an automatic text segment selection mode. Regardless of the mode, when entering the corresponding selection mode, the user's current gaze point is set as the initial position of the cursor.

[0066] Furthermore, in response to entering each selection mode, the cursor is controlled to be displayed at the user's gaze point in a preset form, wherein the preset forms of the cursor include a first form, an aggregate anchor point, a second form, and a third form.

[0067] An embodiment of the present application provides a method for initializing a cursor position. Before selecting a target text block, the cursor is moved within the text content based on the displacement of a preset hand position of the user on a first plane defined by a second coordinate axis and a third coordinate axis to select a location for the cursor. In other words, before selecting a target text block, the cursor is first inserted at the starting point of the text block to be selected. The preset hand position includes, but is not limited to, the position of the thumb tip.

[0068] In one embodiment, in order to meet the text selection requirements for general scenarios (such as web browsing, document reading, etc.), the gestures in the first gesture A combination type in this application include continuously performing a pinching action and releasing a pinching action using a non-index finger and a thumb. In other words, when it is detected that the user continuously performs a pinching action and releases a pinching action using a non-index finger and a thumb, the single text segment manual selection mode is entered. Exemplarily, in response to the user continuously performing a pinching action and releasing a pinching action using the middle finger and thumb, the single text segment manual selection mode is entered. The execution duration of the pinching action is less than the first execution duration. For example, the first execution duration is 0.3 seconds. For example, the pinching action is the action of pinching the index finger and / or middle finger and the thumb together within a preset time.

[0069] Typically, in VR, pinching and releasing the index finger and thumb triggers a function button. Pinch and drag the pinch points to scroll through pages. These two interactions already occupy the more natural pinch gesture. In this application, pinching the middle finger and thumb is used to activate the text selection function, clearly distinguishing between "triggering" and "swiping." Furthermore, the application's actions are memorable and natural.

[0070] In another embodiment, in order to meet the needs of text selection in scenarios such as user chatting and document editing, Figure 3 As shown, the gesture in the first gesture combination type A includes moving the palm across the virtual keyboard plane in the system. In other words, when the user's palm is detected to have moved across the virtual keyboard plane in the system, the system enters the single text segment manual selection mode. At the same time, the key characters on the virtual keyboard disappear and their transparency decreases, indicating that the user has entered the single text segment manual selection mode. In this application, the virtual keyboard plane is used as the interactive space demarcation to clearly distinguish between the interactive instructions of dragging the blinking cursor and text selection.

[0071] In one embodiment, the present application provides a method for selecting a target text block in single text segment manual selection mode. The second gesture operation includes a pinching action and a pinch-point dragging action. The pinch-point dragging action involves dragging your hand in a certain direction while maintaining the pinched state. The pinch point is the point at the thumb's fingertips when the index finger and / or middle finger are pinched together with the thumb.

[0072] Specifically, in the single text segment manual selection mode, in response to the second gesture operation performed by the user, a target text block is selected in the text content using a corresponding text selection rule, including:

[0073] S410: In response to a pinching action and a pinch point dragging action sequentially performed by the user, a preset morphological cursor is controlled to move on the virtual text page based on the resulting displacement of the pinch point, and a target text block is determined based on the position of the preset morphological cursor. Specifically, the displacement of the preset morphological cursor on the virtual text page is controlled based on the displacement of the pinch point by the user after performing the pinching action.

[0074] Furthermore, to facilitate selection of a target text block, the preset cursor shape exemplarily includes a first cursor shape; the first cursor shape includes, but is not limited to, an I-shaped cursor. The pinch action includes a two-finger pinch action; the second gesture operation also includes a release pinch action. The two-finger pinch action is performed by the user using the index or middle finger and thumb; the three-finger pinch action is performed by the user using the index and middle fingers and thumb.

[0075] In response to a pinching action and a pinch point dragging action sequentially performed by a user, controlling the movement of a preset cursor on a virtual text page according to the resulting pinch point displacement, and determining a selected target text block according to the position of the preset cursor, including:

[0076] S411 : In response to a two-finger pinching action performed by the user, confirm that the current position of the first-form cursor is the starting point of a target text block to be selected.

[0077] In response to the user performing a two-finger pinching action using the index finger or the middle finger and the thumb, the current position of the first-form cursor is confirmed to be the starting point of the target text block to be selected.

[0078] S412, based on the pinch point displacement generated by the user performing a pinch point drag action on the first plane, controlling the movement of the first-form cursor on the virtual text page to determine the range of text to be selected; the first plane is formed by the second coordinate axis and the third coordinate axis in the real space coordinate system.

[0079] For example, the coordinate system of the user operation space is constructed, such as Figure 4 As shown, the first coordinate axis in the real-space coordinate system (the coordinate system constructed within the user's real space) is the Z axis, with the positive direction of the Z axis pointing toward the user's body, and the user facing the negative direction of the Z axis. The second coordinate axis is the X axis, with the positive direction of the X axis pointing to the user's right. The third coordinate axis is the Y axis, with the positive direction of the Y axis pointing upward. The first plane is the plane formed by the X and Y axes, and is therefore also called the XY plane.

[0080] S413 : In response to a pinch-release action performed by the user, the current position of the first-form cursor is used as the end point of the target text block.

[0081] S414: Obtain a selected target text block according to the starting point and the ending point.

[0082] Furthermore, the selected target text block is rendered using a first rendering method, which includes but is not limited to setting a background color of the target text block using a preset color, for example, setting the background color of the selected target text block to pink.

[0083] For example, in a single text segment manual selection mode in a general scenario (such as web browsing, document reading, etc.), the operation method includes the following steps:

[0084] (1) In response to the user continuously performing a pinching action with the middle finger and thumb and releasing the pinching action, and the execution time is less than the first execution time, the control system enters the single text segment selection mode, and an I-shaped cursor is displayed in the virtual text page to assist the user in selecting the target text block. The initial position of the I-shaped cursor is the position of the gaze point of the eyes when the middle finger and thumb are pinched, such as Figure 5 As shown in a in .

[0085] (2) Subsequently, the position of the I-shaped cursor is controlled to change with the position of a preset hand position on a first plane formed by the X-axis and the Y-axis, wherein the preset hand position is the thumb tip. When no target text block is obtained, the position of the I-shaped cursor is uploaded to the system in real time and replaced as the starting point of the target text block.

[0086] (3) After the user places the I-shaped cursor to a suitable position through step (2), if it is detected again that the user performs a pinching action with the index finger, middle finger, and thumb, the I-shaped cursor is selected and the cursor position at the time of the pinching action is used as the starting point of the target text block, such as Figure 5 As shown in b in the figure, the user drags the pinch point on the XY plane to control the I-shaped cursor to move to the appropriate position, and then releases the pinch action. The I-shaped cursor position at the time of releasing the pinch action is used as the end point of the target text block. The target text block includes the characters between the starting point and the end point, such as Figure 5 The pink background in c corresponds to the text content.

[0087] For example, in a single text segment manual selection mode in special scenarios (such as chatting, document editing, etc.), the virtual keyboard plane is used as the interactive space division to distinguish it from the interactive instructions of dragging the blinking cursor and text selection, such as Figure 6 As shown in a, the specific operation method includes the following steps:

[0088] (1) When the user's finger passes over the virtual keyboard plane, the blinking cursor in the text input box will be replaced by an I-shaped cursor, and the key characters on the virtual keyboard disappear and reduce their transparency, which is used to prompt the user that they have entered the single text segment manual selection mode, such as Figure 6 As shown in b.

[0089] (2) The I-shaped cursor moves with the displacement of the preset part of the hand (such as the displacement of the thumb tip).

[0090] (3) When the user pinches the index finger or middle finger with the thumb, the position of the I-shaped cursor is used as the starting point of text selection, and the position of the I-shaped cursor when the pinch is released is used as the end point of text selection, such as Figure 6 As shown in c.

[0091] Furthermore, embodiments of the present application provide a method for selecting a target text block based on aggregate anchor points in single text segment manual selection mode. To more quickly select a target text block, the preset cursor shape includes an aggregate anchor cursor; the aggregate anchor cursor includes a front anchor point and a back anchor point. The pinch gesture includes a three-finger pinch gesture; and the second gesture operation also includes a release gesture.

[0092] In response to a pinching action and a pinch point dragging action sequentially performed by a user, controlling the movement of a preset cursor on a virtual text page according to the resulting pinch point displacement, and determining a selected target text block according to the position of the preset cursor, including:

[0093] S411′: In response to a three-finger pinch gesture performed by the user, the current cursor position is set as the insertion position of the aggregate anchor cursor. The insertion position is also called the initial position. Exemplarily, in response to the user performing a three-finger pinch gesture using the index finger, middle finger, and thumb, the current cursor position is set as the insertion position of the aggregate anchor cursor.

[0094] S412', according to the pinch point displacement generated by the user performing the pinch point dragging action along the first target direction on the first plane, control the front anchor point to move toward the first target direction of the virtual text page; wherein the first plane is composed of the second coordinate axis and the third coordinate axis in the real space coordinate system.

[0095] Exemplarily, the front anchor point is controlled to move toward the first target direction of the virtual text page based on the displacement amount of the user dragging the pinch point toward the preset first target direction on the first plane formed by the second coordinate axis and the third coordinate axis while maintaining the three-finger pinch action; wherein the first target direction on the first plane is the positive direction of the Y axis (upward) or the negative direction of the X axis (left). It can be understood that the first target direction of the virtual text page is the corresponding upward or left direction in the page coordinate system, which is specifically determined according to the positive and negative directions of the X and Y axes of the established page coordinate system.

[0096] S413', based on the displacement of the pinch point generated by the user performing the pinch point dragging action along the second target direction on the first plane, the rear anchor point is controlled to move toward the second target direction on the virtual text page; and a target text block is selected based on the front anchor point and the rear anchor point. Exemplarily, based on the displacement of the pinch point detected when the user drags the pinch point toward the preset second target direction on the first plane while maintaining the three-finger pinch action, the rear anchor point is controlled to move toward the second target direction on the virtual text page; and a target text block is selected based on the front anchor point and the rear anchor point. Wherein, the second target direction on the first plane is the negative direction of the Y axis (downward) or the positive direction of the X axis (rightward). It can be understood that the second target direction of the virtual text page is the corresponding downward or rightward direction in the page coordinate system, which is specifically determined according to the positive and negative directions of the X and Y axes of the established page coordinate system.

[0097] Specifically, (1) in response to the user performing a pinching action with the middle finger and thumb, the control system enters the single text block selection mode. After entering the single text segment selection mode, a prompt cursor may appear or not. The prompt cursor may be an I-shaped cursor. In order to provide better feedback to the user, when the prompt cursor appears, the initial position of the insertion cursor is the position of the eye gaze point at the moment of entering the single text segment selection mode, such as Figure 7 As shown in a in .

[0098] (2) In order to adjust the position of the I-shaped cursor, the position change of the I-shaped cursor is controlled according to the position change of the preset part of the hand on the first plane formed by the X-axis and the Y-axis to achieve the appropriate position.

[0099] (3) After detecting that the user adds the index finger to the pinch point, the aggregate anchor cursor is displayed. The aggregate anchor cursor consists of the front anchor point and the back anchor point, such as Figure 7 As shown in b.

[0100] (4) In response to the user dragging the pinch point to the left or up a certain distance, the front anchor cursor is controlled to move left or up; in response to the user dragging the pinch point to the right or down, the back anchor cursor is controlled to move right or down, such as Figure 7 As shown in c.

[0101] It should be noted that when adjusting the positions of the front and rear anchor points in this application, there is no need to re-execute the three-finger pinching action. In other words, the corresponding pinching point in step (3) can be controlled to move upward or leftward, and the front anchor point will be adjusted; it can be understood that as long as the pinching point is controlled to move downward or rightward, the position of the rear anchor point will be adjusted. Furthermore, this application also includes an anchor point locking adjustment method. Figure 8 As shown, specifically including:

[0102] (1) When the user moves the front anchor point beyond the expected forward position or the rear anchor point to the expected rearward position, the user can continue pinching and turn the palm over to use the backhand mode.

[0103] (2) After the backhand mode is activated, the position of the highlighted anchor point can be changed arbitrarily, including up, down, left, and right directions.

[0104] (3) Reversing back to forehand mode will restore normal interaction.

[0105] (4) When the front anchor cursor position is adjusted to behind the rear anchor cursor position, the front and rear anchor icons will also be swapped.

[0106] In one embodiment, the gesture in the first gesture B combination type includes using a non-index finger and a thumb to perform a first pinching action, and dragging the pinching point along the first coordinate axis toward the user's body for a preset distance. In other words, upon detecting that the user uses a non-index finger and a thumb to drag the pinching point along the first coordinate axis toward the user's body for a set distance threshold while maintaining the first pinching action, the control system enters a multi-text segment manual selection mode. Exemplarily, the first coordinate axis is the Z axis. The present application proposes to use the Z-axis spatial movement distance and movement speed to determine the selection intention of multiple text blocks, so that the user can cumulatively select multiple text blocks in a natural and quick manner.

[0107] Furthermore, an embodiment of the present application provides a method for selecting multiple target text blocks in a multi-text segment manual selection mode. The preset cursor shape includes a second cursor shape; the second gesture operation also includes a segment gesture. Exemplarily, the second cursor shape is a blinking cursor.

[0108] In the multi-text segment manual selection mode, in response to the second gesture operation performed by the user, a target text block is selected in the text content using a corresponding text selection rule, including:

[0109] S420, repeatedly using the segmentation gesture and the single text segment manual selection method to select at least two target text blocks; wherein the single text segment manual selection method is an operation method for selecting any target text block in the single text segment manual selection mode.

[0110] Furthermore, the segmented gestures are all gestures in the first gesture B combination type.

[0111] Repeat the segmentation gesture and single text segment manual selection methods to select at least two target text blocks, including:

[0112] S421, after entering the multi-text segment manual selection mode, in response to all gestures in the first gesture combination type B performed by the user, determining to start selecting the next target text block;

[0113] S422, based on the prompt of the second-form cursor, select the next target text block using a manual selection method in a single text segment;

[0114] S423 , repeatedly performing the segmentation gesture and single text segment manual selection method to select multiple target text blocks until the user's needs are met.

[0115] For example, (1) after the user selects a target text block on the current virtual text page using the single text segment selection method, in response to the user pulling the pinch point back toward the body (positive direction of the Z axis), when the pulling back distance exceeds a certain threshold, the multi-text segment manual selection mode is entered and the text selection linkage is temporarily closed. That is, the selection range of the target text is no longer modified by moving the above-mentioned I-shaped cursor, such as Figure 9 a in Figure 9 As shown in b.

[0116] (2) After entering the multi-text segment manual selection mode, the original cursor, such as the I-shaped cursor or the aggregate anchor cursor, will be converted into a blinking cursor, such as Figure 9 As shown in b.

[0117] (3) According to the displacement of the user's pinch point on the first plane formed by the XY axis, the position of the flashing cursor corresponding to the X and Y axes on the virtual text page is controlled to adjust to the selection area of ​​the next target text block to be selected, such as Figure 9 b and Figure 9 As shown in c.

[0118] (4) After moving the blinking cursor to the appropriate insertion position, drag the pinch point away from the body (negative direction of the Z axis) and move it to the set distance threshold, then start selecting another target text block, such as Figure 9 Then continue to use the single text segment selection method to select the new text block, such as Figure 9 As shown in e.

[0119] (5) It supports repeating steps (1) to (4), and the selection state of the new target text block will not affect the selection state of the previously selected target text block.

[0120] In one embodiment, the present application also provides a method for selecting a target text block using the automatic text segment selection mode. In the automatic text segment selection mode, in response to a second gesture operation performed by the user, a corresponding text selection rule is used to select the target text block in the text content, including:

[0121] S430 , based on the acquired segmentation level, using a semantic segmentation method to segment the text content into at least one to-be-selected text block, and selecting at least one target text block according to a second gesture operation.

[0122] Furthermore, the second gesture operation includes a pinch-and-drag action. It is understandable that based on the acquired segmentation level, a semantic segmentation method is used to segment the text content into at least one to-be-selected text block, and at least one target text block is selected according to the second gesture operation, including:

[0123] S431, in the text segment automatic selection mode, determines to enter the corresponding semantic segmentation mode according to the direction in which the user performs the pinch point dragging action on the second coordinate axis. Exemplarily, in response to the user dragging the pinch point along the first direction or the second direction of the second coordinate axis, the corresponding semantic segmentation mode is selected according to the direction of dragging the pinch point; exemplarily, the first direction is the negative direction of the X-axis (left), and the second direction is the positive direction of the X-axis (right). The types of semantic segmentation modes include mechanical semantic segmentation mode and intelligent semantic segmentation mode. Specifically, the user can select the semantic segmentation mode by moving left (negative X direction) or right (positive X direction). For example, moving left is the mechanical semantic segmentation mode, and moving right is the intelligent semantic segmentation mode.

[0124] S432, after determining the semantic segmentation mode, determine the segmentation level according to the pinch point displacement generated by the user performing the pinch point dragging action along the first coordinate axis in the real space coordinate system toward the user's body, and perform segmentation based on the segmentation level to obtain multiple sections of text blocks to be selected; wherein the pinch point displacement is positively correlated with the segmentation level; the types of semantic segmentation modes include mechanical semantic segmentation mode and intelligent semantic segmentation mode.

[0125] Automatic text segment selection modes include single-segment automatic selection mode and multiple-segment automatic selection mode. In the first gesture combination type C, if the gesture includes performing a pinching motion with the non-index finger and thumb, and dragging the pinch point a preset distance along the second coordinate axis or the third coordinate axis, the system is confirmed to enter single-segment automatic selection mode. In other words, in response to the user using the non-index finger and thumb, while maintaining the pinching motion, to drag the pinch point a preset distance along the second coordinate axis or the third coordinate axis, the system is confirmed to enter single-segment automatic selection mode. Exemplarily, the second coordinate axis is the X-axis, and the third coordinate axis is the Y-axis.

[0126] In the first gesture combination type D, the system enters the multi-text segment automatic selection mode when the user performs a pinching motion with the non-index finger and thumb, and drags the pinch point a preset distance along the first coordinate axis. In response to the user maintaining the pinching motion with the non-index finger and thumb, and dragging the pinch point a preset distance along the first coordinate axis, the system enters the multi-text segment automatic selection mode. The first coordinate axis is the X-axis.

[0127] Exemplarily, after detecting a user pinching their middle finger and thumb, the system determines whether to enter single-segment automatic selection mode or multiple-segment automatic selection mode based on how far the user drags the pinch point along the X or Y axis. Specifically, if the user drags the pinch point upwards for more than a certain distance along the X or Y axis, the system enters single-segment automatic selection mode; if the user pulls the pinch point back toward their body (moves it in the positive direction of the Z axis) for a certain distance, the system enters multiple-segment automatic selection mode.

[0128] Furthermore, the semantic segmentation mode is selected based on whether the user drags the pinch point leftward (negative x-axis direction) or rightward (positive x-axis direction). Specifically, if the user drags the pinch point leftward, the mechanical semantic segmentation mode is used (i.e., segmentation is based on characters, words, sentences, and paragraphs); if the user drags the pinch point rightward, the intelligent semantic segmentation mode is used (i.e., segmentation method and level are determined by the AI ​​model).

[0129] After determining the semantic segmentation mode to be used, the segmentation level is determined based on the distance that the user continues to pull the pinch point back toward the body. Figure 10 As shown, taking mechanical semantic segmentation as an example, the larger the Z-axis value of the pinch point, the higher the segmentation level (from individual words to paragraphs).

[0130] Furthermore, the displacement of the preset part of the hand (such as the thumb position or the index finger position) is used to control the movement of the circular cursor on the X and Y axes on the virtual text page. When the circular cursor is in the target text block, the target text block can be selected by pinching, and the target text block can be deselected by pinching again, such as Figure 11 As shown in a in .

[0131] In one embodiment, the present application also provides a method for fine-tuning the range of a selected target text block. In order to more accurately adjust the range of the selected target text block, the present application also provides a backhand fine-tuning mode. After selecting the target text block, the method also includes:

[0132] S510: In response to detecting that the angle between the user's palm and the first coordinate axis is greater than a first preset angle, entering a backhand fine-tuning mode.

[0133] S520: In response to detecting that the angle between the user's palm and the first coordinate axis is smaller than a second preset angle, exit the backhand fine-tuning mode.

[0134] S530: In the backhand fine-tuning mode, insert a front anchor point and a back anchor point according to the start point and the end point of the selected target text block.

[0135] S540: Select a target anchor point according to the pinching action or gaze point performed by the user.

[0136] S550: Adjust the position of the target anchor point according to the displacement of the pinch point generated by the user performing the pinch point dragging action on the first plane.

[0137] For example, after selecting a target text block in a virtual page, the user can make a secondary fine-tuning of the selected text range without having to reselect the target text. The specific methods for secondary fine-tuning include:

[0138] (1) If the palm is flat, Figure 12a As shown, the angle between the positive direction of the Z axis (such as Figure 12b If the angle is greater than a certain angle (such as 110 degrees), it is determined to have entered the backhand fine-tuning mode.

[0139] (2) If the palm plane is less than a certain angle (such as 90 degrees), it is judged as a forehand, and the backhand fine-tuning mode is exited or kept closed.

[0140] (3) After entering the backhand fine-tuning mode, the system will define the positions of the front anchor point and the back anchor point cursor according to the starting point and ending point of the currently selected target text block, and present them to the user visually (such as Figure 12c shown).

[0141] (4) In order to change the selection range more conveniently, this application provides two ways to fine-tune the front and back anchor point positions: manual anchor point selection mode and eye movement anchor point selection mode.

[0142] The first one: Manually select the anchor point mode.

[0143] (1) In the backhand fine-tuning mode, in response to the user performing a pinching action with the index finger, middle finger, and thumb, and controlling the pinching point to move leftward or upward on the first plane beyond a certain threshold (e.g., 10 mm), the system will determine that the user intends to change the front anchor point and highlight the front anchor point, such as Figure 13 As shown in a in .

[0144] (2) In the backhand fine-tuning mode, in response to the user performing a pinching action with the index finger, middle finger, and thumb, and controlling the pinching point to move rightward or downward on the first plane beyond a certain threshold (e.g., 10 mm), the system will determine that the user intends to change the rear anchor point and highlight the rear anchor point, such as Figure 13 As shown in b.

[0145] It should be noted that after pinching the index finger or middle finger with the thumb to select an anchor point, you can only control the displacement of the selected anchor point before releasing and pinching again.

[0146] The second type: eye movement selection anchor point mode.

[0147] In the backhand fine-tuning mode, in response to the user performing a pinching action with the index finger, middle finger, and thumb, the selected anchor point is determined based on the distance between the eye gaze point and the front anchor point and the back anchor point, and the selected anchor point is highlighted, such as Figure 14 As shown in FIG. 1 , the anchor point closest to the gaze point is determined as the selected anchor point, and the anchor point closer to the selected anchor point is highlighted.

[0148] Once you've selected an anchor point, simply continue pinching and moving the pinch point to change its position. These two anchor point selection methods don't conflict with each other and can be combined as needed.

[0149] Among them, the manual anchor point selection mode has higher operation accuracy, and can accurately judge which anchor point the user wants to change when the relative distance between the front and rear anchor points is closer. The eye movement mode is more intuitive, but the eye tracking accuracy of the current device is insufficient, and it is difficult to accurately judge when the two anchor points are close. It can play the greatest advantage when the relative distance between the front and rear anchor points is farther. Therefore, the system can call a better anchor point selection mode by judging the distance between the two anchor points. Specifically, when the distance between the front and rear anchor points is greater than 1.5 times the eye movement error, the eye movement selection anchor point mode is preferred, and the manual anchor point selection mode is also supported. Otherwise, only the manual anchor point selection mode is supported. The eye movement error needs to be determined in combination with the specific head-mounted display and other factors (such as page text line spacing, text size, etc.).

[0150] In one implementation, to increase flexibility, the present application also provides a method for deselecting a target text block.

[0151] That is, in a case where a target text block is selected, in response to a pinching action performed by the user, the selection of the target text block is confirmed to be canceled; in other words, in response to the user performing a pinching action using the index finger or the middle finger and the thumb, the selection of the target text block is confirmed to be canceled;

[0152] When at least two target text blocks are selected, in response to a pinching action performed by the user, the target text block where the cursor is located is determined to be deselected. In other words, in response to the user performing a pinching action using the index finger, middle finger, and thumb, the target text block where the cursor is located is determined to be deselected.

[0153] In one embodiment, the present application also provides a method for selecting multiple text blocks at one time. The preset cursor shape further includes a third cursor shape.

[0154] Based on the obtained segmentation level, a semantic segmentation method is used to segment the text content into at least one segment of text blocks to be selected, including:

[0155] In the case where there are multiple text blocks to be selected, in response to a two-finger pinching action performed by the user, selecting the text block to be selected where the third-form cursor is located;

[0156] After a block of text to be selected is selected, the third-form cursor is controlled to move on the virtual text page based on the number of times the user progressively pulls the pinch point back along the first coordinate axis toward the body and the displacement of each pullback. A movement trajectory of the third-form cursor is generated, and the entire block of text to be selected covered by the movement trajectory is used as the target selected text area.

[0157] In the case where there is a target selected text area, in response to the user dragging the pinch point away from the body along the first coordinate axis, each to-be-selected text block in the target selected text area is determined to be selected.

[0158] Furthermore, the second gesture operation includes a pinch-point drag action; understandably, generating a movement trajectory of the third-form cursor includes:

[0159] During the process of moving the pinch point to the next block of text to be selected, a fixed third-form cursor is generated on the current block of text to be selected, resulting in a first cursor graphic. Starting from the first cursor graphic, a plurality of first cursor graphics are generated, each with a size inversely proportional to the displacement of the pinch point, until the next block of text to be selected. These multiple first cursor graphics constitute the movement trajectory from the current block of text to the next block of text to be selected. It is understood that when the pinch point moves to the next block of text to be selected, the next block of text to be selected becomes the current block of text to be selected, and the above method is similarly executed to generate a fixed third-form cursor.

[0160] For the automatic text segment selection mode, exemplarily, the preset cursor shape is a third cursor shape, which is a circular cursor. The method for selecting multiple text segments to be selected at one time specifically includes the following steps:

[0161] (1) When the middle finger and thumb are pinched together, if the user moves the pinch point along the X-axis or Y-axis and the movement exceeds a certain distance, the system is confirmed to enter the single text segment automatic selection mode; if the user retracts the pinch point toward the body (moves in the positive direction of the Z-axis) for a certain distance, the system is confirmed to enter the multiple text segment automatic selection mode.

[0162] (2) The finger position (such as the thumb position or the index finger position) can control the circular cursor to move on the X-axis and the Y-axis to move to the target text block to be selected. When the user performs a pinching action, the target text block to be selected can be selected. In response to the user performing a pinching action again, the target text block to be selected can be deselected. Figure 11 As shown in a in .

[0163] In one embodiment, a "dragonfly touching the water" method can be used to continuously select multiple text blocks to be selected. That is, after moving the circular cursor to the first text block to be selected that you want to select, pinch your middle finger and thumb. Then keep pinching and retract the pinch point toward the body. The size of the circular cursor will become smaller as the Z-axis value of the pinch point increases. At the same time, the XY position of the circular cursor on the virtual text page is continuously controlled according to the XY position of the pinch point. After the circular cursor is placed at the appropriate position of the text block to be selected, move the pinch point toward the negative half axis of the Z axis. After reaching a certain distance, multiple text blocks to be selected that are covered by the movement track can be selected, such as Figure 11 The above steps can be repeated multiple times to select multiple text blocks to be selected.

[0164] In one embodiment, the present application also provides a method for splitting a block of text to be selected. The preset cursor shape includes a third cursor shape; exemplary, the third cursor shape includes but is not limited to a circular cursor. The second gesture operation includes a pinch point drag action;

[0165] When the text included in a block of text to be selected exceeds a preset threshold, in response to a pinch action performed by the user, controlling the third-modality cursor to move to a target segmentation position according to a pinch point displacement amount generated by the user performing a pinch point drag action on the first plane, wherein the first plane is formed by a second coordinate axis and a third coordinate axis in a real-space coordinate system;

[0166] Hold the pinch point at the target split position for a set time threshold to convert the third form cursor into a split cursor;

[0167] In response to the three-finger pinching action performed by the user, the to-be-selected text block is split into two to-be-selected text sub-blocks at the target split position, such as Figure 15 shown.

[0168] Then, by controlling the pinch point to move toward the first target direction of the first plane, the previous sub-block of text to be selected is selected, and by controlling the pinch point to move toward the third target direction of the first plane, the next sub-block of text to be selected is selected. The third target direction of the first plane can be the upper left direction (the angle between the negative direction of the X axis and the positive direction of the Y axis), such as Figure 15 shown.

[0169] This application can be interactive and easy to remember. Text selection can be achieved simply by pinching and dragging the pinch points, reducing user learning costs and interaction load, improving operational efficiency and accuracy, supporting fast and intuitive text selection area fine-tuning, realizing the "one-stop" text selection function, and supporting cumulative text selection.

[0170] On the first hand, the present application fills the gap in the interaction of text selection in interactive scenarios such as web browsing. In the existing virtual environment of web browsing interaction, "pinching the index finger and thumb and releasing" triggers a function key, and "pinching the index finger and thumb and dragging" slides the page. The above interactions occupy the main interaction mode of "pinching", and text selection is not supported in this scenario. The present application proposes an interaction method based on the "middle finger and thumb" pinching to activate text selection and subsequent text selection. In addition, the present application uses parameters such as the movement and distance of the user's pinch point in the Z-axis direction for "cumulative text selection" (multi-text block selection), which fills the interaction gap in this type of scenario.

[0171] (1) Reduce user learning costs: Since the middle finger is close to the index finger, users can quickly master the pinching method with the middle finger after becoming familiar with pinching with the index finger and thumb.

[0172] (2) Reduce interaction load: By pinching different fingers to clearly point to different interactions, the user's cognitive load and the amount of information they need to think about during the interaction are reduced. Compared to two quick pinches to activate text selection mode, this application only requires a single pinch of the middle finger and thumb to activate it, which reduces the user's physical load. At the same time, Z-axis movement (represented by pulling back or extending toward the body) to activate cumulative text selection is also quite natural and intuitive.

[0173] Secondly, optimize text selection interactions in scenarios such as text input.

[0174] Text input scenarios in existing virtual environments (such as Word document editing, notepad, and document reading and annotation) generally include two interface elements: a virtual keyboard and a virtual input interface, along with various interactive tasks such as text entry, moving the input position (moving the blinking cursor), and text selection. Text entry is primarily performed on the virtual keyboard; moving the blinking cursor position is achieved by pinching and dragging with the index finger and thumb; text selection mode is activated using a double-quick pinch. This application proposes a design that uses the keyboard plane as the interface for dividing the interactive space.

[0175] (1) Reduce interaction load: Compared to double-pinch to call up text selection mode, reaching across the keyboard surface can be more convenient and intuitive. In addition, the operation surface of the virtual keyboard is usually fixed. When the user is operating blindly (only looking at the text interface above), they can also complete the operation of reaching across the keyboard surface by roughly estimating the distance.

[0176] (2) Improve operational efficiency and accuracy: The existing double-pinch to call out text selection mode is based on the phrase closest to the flashing cursor after double-pinch, and places front and back anchor cursors before and after the phrase. The user needs to use their eyes to look at a certain anchor point a second time and pinch and drag the anchor point again, constantly adjusting the front and back anchor points to correct the selected content, which greatly affects the interaction efficiency and causes a decline in user experience. In the solution proposed in this application, an I-shaped cursor appears after the palm passes over the keyboard plane and moves with the movement of the hand. It is only necessary to move the I-shaped cursor to the appropriate position, pinch and drag, and then release it to select the content you intend to select. If there is no operational error or change of intention, no repeated adjustment is required.

[0177] Thirdly, it supports quick and intuitive fine-tuning of text selection areas.

[0178] The existing interaction for adjusting the selection area usually requires using the eyes to stare at the front and rear anchor cursors and then pinching the fingers to drag the cursor. This application proposes a method that uses the initial displacement direction of the pinch point after backhand pinching to determine whether the user wants to change the front anchor point or the rear anchor point.

[0179] (1) Improved operation accuracy: In existing interactive interfaces, the cursors of the front and rear anchor points are usually small (generally close to the size of text), which poses a challenge to the accuracy of eye tracking (especially when the front and rear anchor points are very close). In addition, the eye movement modality is more unstable than other modalities such as hand movement, head movement, and body movement (that is, the position of a person's eye gaze point is more likely to change quickly), which poses a challenge to the user's hand-eye coordination and grasp of the pinching timing. It is easy for the eyes to accidentally look at other positions at the moment of pinching, resulting in a failure in anchor point selection. In this application, the movement distance and direction of the front and rear anchor points are used to determine the movement of the front and rear anchor points in the initial period after the user's finger is pinched. The existing manual tracking system is more accurate, and the system will decide to change the position of the front or rear anchor point only after determining that the displacement of the pinching point exceeds a certain distance, further improving the judgment accuracy.

[0180] (2) Improve the convenience and continuity of interaction: In this solution, users do not need to follow the steps to select a specific anchor point and then pinch and adjust the displacement. They only need to pinch and move in a certain direction to include the anchor point selection and directly move the anchor point position.

[0181] Fourthly, it supports "one-stop" text selection.

[0182] The text selection ideas of the existing scheme and the above schemes are all based on moving from one end of the final intended framed content to the other end to form a framed selection range. One end of the text selection content has been determined during the initial kneading. If you want to change the position of that end, you need to make subsequent adjustments, and the subsequent adjustments involve repeatedly kneading the front and back anchor points and dragging them. This application proposes a text selection method based on continuous kneading that completes text selection and fine-tuning in one go.

[0183] (1) Improve interaction efficiency: After pinching the middle finger and thumb together, the user can call out the I-shaped cursor and move it to a relatively accurate position (it only needs to be within the text area to be selected, not necessarily at one end). This can help users reduce the decision-making cost and interaction accuracy requirements at the beginning of the interaction. Then, simply add the index finger to the pinch to call out the aggregate anchor cursor. At this time, the movement of the pinch point directly controls the selection and movement of the front and back anchor points.

[0184] (2) Improved interaction continuity: In this solution, you do not need to release the three-finger pinch to fine-tune the front and back anchor points. Simply flip your palm to the backhand to lock and adjust the currently highlighted anchor cursor, and it supports full-scale movement up, down, left, and right. After the adjustment is completed, flip your palm to the forehand to unlock a certain anchor point. Finally, after the selection or adjustment is completed, release the pinch to obtain a precise text selection area, which is convenient for subsequent copying, pasting, or other operations.

[0185] Figure 16 A schematic diagram of the structure of a text selection device according to an embodiment of the present application is shown. Exemplarily, the text selection device includes: a display module 210 , an operation type recognition module 220 , a mode entry module 230 and a text selection module 240 .

[0186] Display module 210, for displaying text content on a virtual text page in a virtual reality scene via a head-mounted display device;

[0187] The operation type identification module 220 is configured to identify the type of gesture operation performed by the user in response to a first gesture operation performed by the user;

[0188] Mode entry module 230, for entering a corresponding text selection mode according to the gesture operation type; the types of text selection modes include single text segment manual selection mode, multiple text segment manual selection mode, and text segment automatic selection mode;

[0189] The text selection module 240 is configured to select a target text block in the text content using a corresponding text selection rule in response to a second gesture operation performed by the user in a corresponding text selection mode.

[0190] It can be understood that the device of this embodiment corresponds to the text selection method of the above embodiment, and the optional items in the above embodiment are also applicable to this embodiment, so they will not be described again here.

[0191] The present application also provides a terminal device. Exemplarily, the terminal device includes a processor and a memory, wherein the memory stores a computer program, and the processor runs the computer program to enable the terminal device to execute the functions of the various modules in the above-mentioned text selection method or the above-mentioned text selection method device.

[0192] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0193] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving an execution instruction.

[0194] This application also provides a computer-readable storage medium for storing the computer program used in the terminal device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0195] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A text selection method, characterized in that: Applicable to an extended reality system, the method includes: Displaying text content on a virtual text page in a virtual reality scene through a head-mounted display device; In response to a first gesture operation performed by a user, identifying a type of gesture operation performed by the user; Entering a corresponding text selection mode according to the gesture operation type, wherein the text selection modes include a single text segment manual selection mode, a multiple text segment manual selection mode, and an automatic text segment selection mode; In the corresponding text selection mode, in response to the second gesture operation performed by the user, selecting a target text block in the text content using a corresponding text selection rule; Wherein, the second gesture operation includes a pinching action and a pinch point dragging action; In response to the second gesture operation performed by the user, selecting a target text block in the text content using a corresponding text selection rule includes any one of the following: Item 1: In the single text segment manual selection mode, in response to the pinching action and the pinch point dragging action sequentially performed by the user, a movement amount of a preset cursor on the virtual text page is controlled according to a resulting displacement of the pinch point, and a selected target text block is determined according to a position of the preset cursor; Item 2: The preset shape cursor includes a second shape cursor; the second gesture operation also includes a segmented gesture; In the multi-text segment manual selection mode, repeatedly using the segmentation gesture and the single text segment manual selection method to select at least two target text blocks; wherein the single text segment manual selection method is an operation method for selecting any target text block in the single text segment manual selection mode; Item 3: In the text segment automatic selection mode, based on the acquired segmentation level, a semantic segmentation method is used to segment the text content into at least one to-be-selected text block, and at least one target text block is selected according to the second gesture operation.

2. The text selection method according to claim 1, wherein: The preset shape cursor includes a first shape cursor; the pinching action includes a two-finger pinching action; the second gesture operation also includes a release pinching action; In response to the pinching action and the pinch point dragging action sequentially performed by the user, controlling the movement of a preset cursor on the virtual text page according to a generated pinch point displacement, and determining a selected target text block according to a position of the preset cursor, including: In response to the two-finger pinching action performed by the user, confirming that the current position of the first-shaped cursor is the starting point of a target text block to be selected; controlling a movement amount of the first-modality cursor on the virtual text page to determine a range of text to be selected based on a displacement amount of the pinch point generated by the user performing the pinch-point dragging action on the first plane, wherein the first plane is formed by a second coordinate axis and a third coordinate axis in a real-space coordinate system; In response to a pinch-release action performed by the user, taking the current position of the first-shaped cursor as an end point of the target text block; A selected target text block is obtained according to the starting point and the ending point.

3. The text selection method according to claim 1, wherein: The preset shape cursor includes an aggregate anchor cursor; the aggregate anchor cursor includes a front anchor point and a back anchor point; the pinch action includes a three-finger pinch action; the second gesture operation also includes a release pinch action; In response to the pinching action and the pinch point dragging action sequentially performed by the user, controlling the movement of a preset cursor on the virtual text page according to a generated pinch point displacement, and determining a selected target text block according to a position of the preset cursor, including: In response to the three-finger pinching action performed by the user, setting the current position of the cursor as the insertion position of the aggregate anchor cursor; controlling the front anchor point to move toward the first target direction of the virtual text page based on a pinch point displacement caused by the user performing a pinch point dragging action along a first target direction on a first plane, wherein the first plane is formed by a second coordinate axis and a third coordinate axis in a real space coordinate system; controlling the rear anchor point to move toward the second target direction on the virtual text page according to a displacement of the pinch point generated by the user performing the pinch point dragging action along the second target direction on the first plane; A target text block is selected according to the front anchor point and the back anchor point.

4. The text selection method according to claim 1, wherein: The segmented gestures are all gestures in the first gesture combination type B; the repeatedly using the segmented gestures and the single text segment manual selection method to select at least two target text blocks includes: After entering the multi-text segment manual selection mode, determining to start selecting the next target text block in response to all gestures in the first gesture combination type B performed by the user; Based on the prompt of the second-form cursor, the next target text block is selected using the manual selection method in the single text segment; The segmentation gesture and the single text segment manual selection method are repeatedly performed to select multiple target text blocks until user needs are met.

5. The text selection method according to claim 1, wherein: The step of segmenting the text content into at least one to-be-selected text block using a semantic segmentation method based on the acquired segmentation level, and selecting at least one target text block according to the second gesture operation, includes any one of the following: Item 1: In the automatic text segment selection mode, determining to enter a corresponding semantic segmentation mode according to the direction in which the user performs the pinch point drag action on the second coordinate axis; wherein the types of semantic segmentation modes include mechanical semantic segmentation mode and intelligent semantic segmentation mode; After determining the semantic segmentation mode, determining the segmentation level based on a pinch point displacement amount generated by the user performing the pinch point dragging action along a first coordinate axis in the real space coordinate system toward the user's body, and performing segmentation based on the segmentation level to obtain a plurality of to-be-selected text blocks; wherein the pinch point displacement amount is positively correlated with the segmentation level; Item 2: The preset shape cursor also includes a third shape cursor; In the case where there are multiple text blocks to be selected, in response to the two-finger pinching action performed by the user, selecting the text block to be selected where the third-form cursor is located; After a section of the to-be-selected text block is selected, the third-form cursor is controlled to move on the virtual text page based on the number of times the user progressively pulls the pinch point back along the first coordinate axis toward the body and the displacement of each pullback, and a movement trajectory of the third-form cursor is generated, with the entire at least one section of the to-be-selected text block covered by the movement trajectory being used as a target selected text area; In the case where the target selected text area exists, in response to the user dragging the pinch point away from the body along the first coordinate axis, each to-be-selected text block in the target selected text area is determined to be selected.

6. The text selection method according to claim 5, characterized in that: Generating the movement trajectory of the third-form cursor includes: In the process of moving the pinch point to the next section of the text block to be selected, a fixed third-form cursor is generated on the current text block to be selected to obtain a first cursor graphic; starting from the first cursor graphic, a plurality of first cursor graphics with sizes inversely proportional to the displacement of the pinch point are generated until the next section of the text block to be selected, and multiple first cursor graphics constitute the movement trajectory from the current text block to the next section of the text block to be selected.

7. The text selection method according to claim 5, characterized in that: When the text included in a section of the to-be-selected text block exceeds a preset threshold, in response to the pinching action performed by the user, controlling the third-modality cursor to move to a target segmentation position according to a pinch point displacement generated by the user performing the pinch point dragging action on a first plane, wherein the first plane is formed by a second coordinate axis and a third coordinate axis in a real-space coordinate system; The pinch point is kept at the target split position for a set time threshold, and the third-form cursor is converted into a split cursor; In response to the three-finger pinching action performed by the user, the to-be-selected text block is split into two to-be-selected text sub-blocks at the target split position.

8. The text selection method according to claim 1, wherein: The method further comprises at least one of the following four items: Item 1: After selecting the target text block, the method further includes: In response to detecting that the angle between the user's palm and the first coordinate axis is greater than a first preset angle, entering a backhand fine-tuning mode; In response to detecting that the angle between the palm of the user and the first coordinate axis is smaller than a second preset angle, exiting the backhand fine-tuning mode; In the backhand fine-tuning mode, inserting a front anchor point and a back anchor point according to the starting point and the ending point of the selected target text block; selecting a target anchor point based on a pinching action or a gaze point performed by the user; adjusting a position of the target anchor point according to a pinch point displacement amount generated by the user performing a pinch point drag action on a first plane, where the first plane is formed by a second coordinate axis and a third coordinate axis in a real space coordinate system; Item 2: When entering the corresponding selection mode, setting the acquired current gaze point of the user as the initial position of the cursor; Before selecting the target text block, controlling the movement of the cursor within the text content according to the displacement of the preset position of the user's hand on the first plane formed by the second coordinate axis and the third coordinate axis to select a position for placing the cursor; Item 3: In a case where the target text block is selected, in response to a pinching action performed by the user, confirming cancellation of the selection of the target text block; In a case where at least two target text blocks are selected, in response to the pinching action performed by the user, determining to cancel the selection of the target text block where the cursor is located; Item 4: The text segment automatic selection mode includes a single text segment automatic selection mode and a multiple text segment automatic selection mode; entering the corresponding text selection mode according to the gesture operation type includes: If the gesture operation type is a first gesture A combination type, controlling the system to enter a single text segment manual selection mode; the gestures in the first gesture A combination type include continuously performing a pinch action, releasing a pinch action, or using a palm to pass through the virtual keyboard plane in the system; If the gesture operation type is a first gesture combination type B, controlling the system to enter a multi-text segment manual selection mode; the gesture in the first gesture combination type B includes performing a pinch action and dragging the pinch point a preset distance along the first coordinate axis toward the user's body; If the gesture operation type is the first gesture C combination type or the first gesture D combination type, the system is controlled to enter the text segment automatic selection mode; the gesture in the first gesture C combination type includes performing the pinching action and dragging the pinch point a preset distance along the second coordinate axis or the third coordinate axis; the gesture in the first gesture D combination type includes performing the pinching action and dragging the pinch point a preset distance along the first coordinate axis.

9. An extended reality system, characterized in that: include: Head-mounted display equipment, camera equipment and processing center; The head mounted display device is used to display text content on a virtual text page in a virtual reality scene; The camera device is used to take a picture of the user's hand to obtain a gesture operation image; The processing center is configured to recognize a first gesture operation and / or a second gesture operation of a user based on the gesture operation image, so as to implement the text selection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text segment selecting method and field selecting method, device and terminal

    CN103186345A

  • Gesture movement in artificial reality environment

    CN117590931A