Information interaction method and device, electronic equipment and storage medium

The technology of acquiring and processing images through wearable devices and using large models to identify and output continuation content is solved by solving the problem of user interruptions in reading and improving the continuity and user experience of reading.

CN120371136APending Publication Date: 2025-07-25SHANGHAI XIAODU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510529770.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Users are prone to interruption of reading due to eye fatigue or inconvenience to handheld devices during reading, resulting in poor reading experience.

Method used

The image of the target content is obtained through the wearable device, and the continuation content is identified and determined using a large model, which is converted into voice output to meet the user's continuation reading needs.

Benefits of technology

When users are inconvenient to continue reading, providing continuous content through wearable devices improves the continuity and user experience of reading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371136A_ABST
    Figure CN120371136A_ABST
Patent Text Reader

Abstract

The invention provides an information interaction method, and relates to the technical field of artificial intelligence, in particular to the technical field of large model and data processing. According to the specific implementation scheme, in response to the situation that the intention of a target object is determined to obtain continuous content of target content according to input information of the target object, a target image containing the target content is obtained through wearable equipment; determining a target resource corresponding to the target content according to the target image; determining continuous content of the target content according to the target resource; and transmitting the continuation content to the wearable device so as to output the continuation content via the wearable device. The invention further provides an information interaction device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, and more particularly to the fields of large models and data processing technologies. More specifically, the present disclosure provides an information interaction method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] When users read articles or news, situations often occur where it is inconvenient to read or the reading is interrupted, resulting in fragmented reading for users and a poor reading experience. Summary of the Invention

[0003] The present disclosure provides an information interaction method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to a first aspect, there is provided an information interaction method, the method comprising: in response to determining, based on input information of a target object, that the intention of the target object is to obtain a continuation of target content, acquiring, by using a wearable device, a target image including the target content; determining, based on the target image, a target resource corresponding to the target content; determining, based on the target resource, a continuation of the target content; and sending the continuation to the wearable device for outputting the continuation via the wearable device.

[0005] According to a second aspect, there is provided an information interaction apparatus, the apparatus comprising: an image acquisition module configured to, in response to determining, based on input information of a target object, that the intention of the target object is to obtain a continuation of target content, acquire, by using a wearable device, a target image including the target content; a resource determination module configured to determine, based on the target image, a target resource corresponding to the target content; a content determination module configured to determine, based on the target resource, a continuation of the target content; and an output module configured to send the continuation to the wearable device for outputting the continuation via the wearable device.

[0006] According to a third aspect, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, enable the at least one processor to execute the method provided by the present disclosure.

[0007] According to a fourth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided by the present disclosure.

[0008] According to a fifth aspect, there is provided a computer program product including a computer program stored on at least one of a readable storage medium and an electronic device, the computer program, when executed by a processor, implementing the method provided according to the present disclosure.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which an information interaction method and apparatus according to an embodiment of the present disclosure can be applied;

[0012] Figure 2 is a flowchart of an information interaction method according to an embodiment of the present disclosure;

[0013] Figure 3 is a flowchart of an information interaction method according to another embodiment of the present disclosure;

[0014] Figure 4 is a block diagram of an information interaction apparatus according to an embodiment of the present disclosure; and

[0015] Figure 5 is a block diagram of an electronic device for an information interaction method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0017] With the development of technology, wearable intelligent devices are becoming increasingly popular, such as smart glasses. Users can take pictures using the camera on the smart glasses and can interact by combining the capabilities of large models. Currently, smart glasses on the market can identify objects in the picture after taking a photo and then combine the large model to give information about the picture. However, they cannot meet more interaction needs of users.

[0018] For example, in some scenarios, when users are taking a bus or subway, they will use their free time to read books or news. When their eyes get tired, they get off the vehicle, or it is inconvenient to hold a book, mobile phone, or tablet, reading will be interrupted.

[0019] Embodiments of the present disclosure address this scenario and use wearable smart devices to provide users with the interactive need for continuous reading.

[0020] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0021] In the technical solutions of the present disclosure, before obtaining or collecting the user's personal information, the user's authorization or consent has been obtained.

[0022] Figure 1 It is a schematic diagram of an exemplary system architecture to which the information interaction method and device can be applied according to an embodiment of the present disclosure. It should be noted that Figure 1 The illustration shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.

[0023] As Figure 1 shown, the system architecture 100 according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for the communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired and / or wireless communication links, etc.

[0024] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. The terminal device 101 may be various electronic devices, including but not limited to wearable smart devices such as smart glasses and smart watches, head-mounted terminal devices, etc.

[0025] The server 103 may be a server that provides various services, such as a background management server (only an example) that provides support for the information obtained by the user using the terminal device 101. The background management server may analyze and process data such as the received user requests, and send feedback information such as audio data and image data of the user requests to the terminal device 101.

[0026] The information interaction method provided by the embodiments of the present disclosure can generally be executed by the server 103. Correspondingly, the information interaction device provided by the embodiments of the present disclosure can generally be set in the server 103.

[0027] Figure 2 It is a flowchart of an information interaction method according to an embodiment of the present disclosure.

[0028] like Figure 2 As shown, the information interaction method 200 includes operations S210 to S240.

[0029] In operation S210, in response to determining, according to the input information of the target object, that the target object intends to obtain continuation content of the target content, a target image including the target content is obtained by using a wearable device.

[0030] In the embodiments of the present disclosure, the target object may refer to a user who is using or holding a wearable device. The wearable device may refer to a terminal device such as smart glasses, smart headphones, and smart watches that has image acquisition and audio broadcast functions. The input information may be operation information initiated by the user through button clicks, gesture control, etc., or may be instruction information input by the user in text or voice form.

[0031] For example, the wearable device includes an interactive button for obtaining the continuation content of the target content, and the user can initiate a request for obtaining the continuation content through the wearable device by clicking the button.

[0032] For example, a user can give a "continue playing subsequent content" command to a wearable device in the form of voice to obtain the continuation of the target content.

[0033] In an embodiment of the present disclosure, the target content may refer to a portion of the content in a handheld book that a user is reading, or a portion of the content in a book, article, news, or video that is viewed through an electronic device.

[0034] In one example, the user may not be able to obtain the subsequent content of the target content or may be interrupted and pause reading. In another example, the user may need to obtain the extended content or interpretation content of the target content and pause the current reading. In these cases, the subsequent content, extended content, and interpretation content of the target content can all be referred to as continued content.

[0035] If the user needs to continue to obtain the continued content, he can express his intention to obtain the continued content by inputting information. After the user's intention to obtain the continued content is identified based on the user's input information, the wearable device with image acquisition function can be used to acquire an image of the target content specified by the target object to obtain a target image containing the target content.

[0036] For example, when a user feels eye fatigue while reading news on a mobile terminal, the user can click a button on the smart glasses to express the intention of obtaining continued content. In response to this intention, the smart glasses can take a picture of the article on the mobile terminal to obtain a target image containing the news content.

[0037] For another example, when a user feels eye fatigue while reading news on a mobile terminal, the user can input voice to the smart glasses, and the smart glasses can recognize from the input voice whether the user has the intention of obtaining continued content. Or the smart glasses can send the input voice to a large model, and the large model can recognize whether the user has the intention of obtaining continued content. When the large model recognizes that the user has the intention of continued reading, the large model can send this intention to the smart glasses. In response to this intention, the smart glasses can take a picture of the article on the mobile terminal to obtain a target image containing the news content.

[0038] In operation S220, according to the target image, determine the target resource corresponding to the target content.

[0039] In an embodiment of the present disclosure, the target resource may refer to a file or a web page containing content such as articles, news, videos, etc. that the user is browsing. For example, the target resource corresponding to the target content can be obtained by retrieving a database, retrieving content in a publicly available web page, etc.

[0040] In operation S230, according to the target resource, determine the continued content of the target content.

[0041] In an embodiment of the present disclosure, the continued content refers to the content that the user wants to continue to obtain by expressing input information.

[0042] For example, after the user stops reading the target content, the user clicks a button. From this, it can be determined that the user's intention indicates a need to obtain subsequent content, and the continued content refers to other content information after the target content.

[0043] For example, after the user stops reading the target content, the user issues an instruction to the wearable device saying "I don't want to read anymore. Summarize the article for me." From this, it can be determined that the user's intention indicates a need to obtain a summary of the article, and the continued content can refer to the content that summarizes the information of the target resource.

[0044] In operation S240, send the continued content to the wearable device so that the continued content can be output via the wearable device.

[0045] For example, the continued content can be converted into voice, and the voice can be output using the audio output link of the wearable device, enabling the user to obtain the continued content in a listening manner.

[0046] According to an embodiment of the present disclosure, in response to determining that the user's intention is to obtain the continuation content of the target content, a wearable device is used to obtain a target image containing the target content, a target resource corresponding to the target content is obtained according to the target image, the continuation content required by the user is obtained from the target resource, and the continuation content is output to the user through the wearable device. In the case where the user is inconvenient to continue obtaining content, the wearable device is used to meet the user's need for obtaining the continuation content, thereby improving the user's interaction experience.

[0047] In an embodiment of the present disclosure, sending the continuation content to the wearable device for outputting the continuation content via the wearable device includes: converting the continuation content into speech; and outputting the speech via the wearable device.

[0048] For example, the continuation content can be converted from text form to speech through TTS (Text To Speech) technology, and the user can obtain the continuation content by playing the speech through the wearable device.

[0049] In an embodiment of the present disclosure, determining the target resource corresponding to the target content according to the target image includes: extracting visual objects from the target image, where the visual objects include at least one of text, illustrations, and video elements; and determining the target resource corresponding to the target content according to the visual objects.

[0050] According to an embodiment of the present disclosure, image recognition can be performed on the target image to extract visual objects such as text information, illustrations, and video elements in the target image. The video elements include, for example, titles, objects in the video frame, progress bars, etc. When the visual object is text, target resources such as novels, news, and papers can be searched according to the corresponding text information; when the visual object is an illustration, target resources such as comics and picture books can be searched according to the illustration; when the visual object is a video element, target resources such as related videos can be searched according to the title, object, etc. in the video element.

[0051] In an embodiment of the present disclosure, determining the target resource corresponding to the target content according to the visual object includes at least one of the following: retrieving the resource library according to the visual object to obtain the target resource corresponding to the target content; and calling a large model to generate a response message containing the target resource information based on the visual object.

[0052] The resource library may refer to a database configured in the server that includes various resources such as papers, news, and videos, or may refer to an open retrieval website, a resource forum, etc. Based on the visual object, the target resource matching the target content can be obtained from the resource library through methods such as text retrieval, image feature matching, and vector retrieval.

[0053] In an embodiment of the present disclosure, information extraction and summarization of visual objects can also be performed through a large model to determine target resources, and response information containing target resource information output by the large model can be obtained.

[0054] For example, if the visual object is text, the large model can be used to understand and retrieve the text to obtain target resources related to the visual object.

[0055] For example, if the visual object is an illustration or a video element, the large model can be used to extract features and match the illustration or video element to retrieve target resources related to the visual object.

[0056] According to an embodiment of the present disclosure, when the user is inconvenient to continue obtaining content, the wearable device can be used to obtain the target content that the user is currently reading or watching, and the large model can be used to extract key information and retrieve the target content to obtain target resources related to the target content, so as to meet the user's need for obtaining continuous content and improve the user interaction experience.

[0057] In an embodiment of the present disclosure, determining the continuous content of the target content according to the target resources includes: determining the sub-intentions of the target object, where the sub-intentions include at least one of obtaining the subsequent content of the target content in the target resources, obtaining the extended content of the target content, and obtaining the interpretation content of the target content; in response to the sub-intentions, generating continuous content based on the target resources, where the continuous content includes at least one of the subsequent content, extended content, and interpretation content of the target content.

[0058] In an embodiment of the present disclosure, the intentions of the target object may be relatively diverse. The sub-intentions may refer to more specific requirements determined according to the input information of the target object after it is obtained that the intention of the target object is to obtain the continuous content of the target content.

[0059] In one embodiment, the sub-intention is expressed as obtaining the subsequent content of the target content in the target resources. In response to the sub-intention of obtaining the subsequent content of the target content in the target resources, determining the progress of the target content in the target resources according to the target image; and obtaining the subsequent content of the target content in the target resources according to the progress.

[0060] The subsequent content may refer to the content that has not been read subsequently when the target object stops reading.

[0061] For example, the input information of the target object is "Continue to read the following content". Based on the input information of the target object, the sub-intention of the target object can be determined as obtaining the subsequent content of the target content. The reading progress of the target object can be determined according to the target image, and then the unread subsequent content can be determined. The subsequent content can be converted into voice and sent to the wearable device, or the subsequent content can be sent to the wearable device, and the wearable device converts the subsequent content into voice. By playing the subsequent content in voice form through the wearable device, the user can obtain the continuation content in a listening manner.

[0062] In another embodiment, the sub-intention can be expressed as obtaining the extended content of the target content. In response to the sub-intention being to obtain the extended content of the target content, the resources associated with the target resource are determined; and based on the target resource and the resources associated with the target resource, the extended content of the target content is generated.

[0063] The extended content may refer to the content information of other resources related to the target content viewed by the target object, such as the references cited in an article, the links inserted in a video, etc.

[0064] For example, the input information of the target object is "I want to see the references cited here". Based on the input information of the target object, the sub-intention of the target object can be determined as obtaining the content information of other articles related to the target content. The corresponding target resource is determined according to the target image, and the references cited at the specified position of the target object are found from the target resource. The references can be converted into voice and sent to the wearable device, or the reference text can be sent to the wearable device, and the wearable device converts the references into voice. By playing the reference text in voice form through the wearable device, the user can obtain the continuation content in a listening manner.

[0065] In another embodiment, the sub-intention can be expressed as obtaining the interpretation content of the target content. In response to the sub-intention being to obtain the interpretation content of the target content, at least one of the summary and interpretation information of the target content is generated according to the context of the target content in the target resource.

[0066] The interpretation content may refer to analysis content such as the summary, conclusion, highlights, etc. of the target content.

[0067] For example, if the input information of the target object is "Help me interpret the news", the sub-intention of the target object can be determined as obtaining in-depth interpretation information of the news according to the input information of the target object. The corresponding news resources can be determined based on the target image, and the news resources can be deeply interpreted from multiple perspectives through a large model to obtain multiple news highlights. The news highlights can be converted into voice and sent to the wearable device, or the news highlights can be sent to the wearable device, and the wearable device converts the news highlights into voice. Playing the news highlights in voice form through the wearable device enables the user to obtain continuous content in a listening manner.

[0068] In the embodiments of the present disclosure, the target content may be at least one of book content, news content, and video content.

[0069] When the target content is book content and news content, the subsequent content, extended content, or interpretation content of the book content and news content can be obtained as continuous content, the continuous content can be converted into voice, and played on the wearable device.

[0070] When the target content is video content, due to the functional limitations of the wearable device, in the case where the subsequent content of the video cannot be played, the audio of the subsequent content of the video content in the target resource can be obtained as continuous content according to the sub-intention, or the summary information of the video content can be generated as continuous content.

[0071] For example, if the input information of the target object is "Obtain video summary", the sub-intention of the target object can be determined as obtaining the summary information of the video according to the input information of the target object. The corresponding video resources can be determined based on the target image, and then the text information such as titles and subtitles in each frame of the video resources can be extracted through a large model, and the content description information of each frame of the image can be extracted, and then the text information and the image content description information can be integrated to obtain video description information, and then the video description information can be summarized to obtain the video summary. The video summary can be converted into voice and sent to the wearable device, or the video summary can be sent to the wearable device, and the wearable device converts the video summary into voice. Playing the video summary in voice form through the wearable device enables the user to obtain continuous content in a listening manner.

[0072] According to the embodiments of the present disclosure, after receiving the user's request for continuous reading, by identifying the sub-intention of the user, the continuous content required by the user can be obtained, which can meet the refined needs of the user and improve the user's interaction experience.

[0073] Figure 3 It is a flowchart of an information interaction method according to another embodiment of the present disclosure.

[0074] As Figure 3As shown, the information interaction method includes operations S301 to S310.

[0075] In operation S301, determine the intention of the target object according to the input information of the target object.

[0076] In operation S302, determine whether the intention of the target object indicates obtaining the continuation content of the target content. If the intention of the target object indicates obtaining the continuation content of the target content, then continue with operation S303. If the intention of the target object has other meanings, then end the operation.

[0077] In operation S303, use the wearable device to obtain the target image containing the target content.

[0078] In operation S304, recognize the target image.

[0079] In operation S305, determine whether a visual object is recognized. If a visual object can be recognized from the target image, then continue with operation S306. If not, then end the operation.

[0080] In operation S306, retrieve the target resource corresponding to the target content according to the visual object.

[0081] For example, the target resource can be retrieved from the resource library based on the visual object, or the large model can be requested to make the large model return the response information containing the target resource information.

[0082] In operation S307, determine whether the target resource is retrieved. If the corresponding target resource is detected, then continue with operation S308. If not detected, then end the operation.

[0083] In operation S308, generate the continuation content according to the target resource.

[0084] For example, according to whether the sub-intention of the user is to obtain subsequent content, extended content, or interpretation content, one of the subsequent content, extended content, or interpretation content can be generated based on the target resource.

[0085] In operation S309, slice the continuation content and convert the sliced content into speech.

[0086] In the embodiments of the present disclosure, since the continuation content may be relatively long and it is difficult to directly convert the continuation content into speech, the continuation content can be sliced to obtain the sliced content. The sliced content is, for example, short sentences. Then the sliced content is converted into speech.

[0087] In operation S310, send the speech to the wearable device.

[0088] According to the embodiments of the present disclosure, the present disclosure also provides an information interaction device.

[0089] Figure 4 It is a block diagram of an information interaction device according to an embodiment of the present disclosure.

[0090] As Figure 4 shown, the information interaction device 400 includes an image acquisition module 410, a resource determination module 420, a content determination module 430, and an output module 440.

[0091] The image acquisition module 410 is configured to, in response to determining that the intention of the target object is to obtain the continuation content of the target content according to the input information of the target object, use the wearable device to acquire a target image including the target content.

[0092] The resource determination module 420 is configured to determine a target resource corresponding to the target content according to the target image.

[0093] The content determination module 430 is configured to determine the continuation content of the target content according to the target resource.

[0094] The output module 440 is configured to send the continuation content to the wearable device so as to output the continuation content via the wearable device.

[0095] According to an embodiment of the present disclosure, the resource determination module 420 includes an extraction sub-module and a resource determination sub-module.

[0096] The extraction sub-module is configured to extract visual objects from the target image, where the visual objects include at least one of text, illustrations, and video elements. The resource determination sub-module is configured to determine a target resource corresponding to the target content according to the visual objects.

[0097] According to an embodiment of the present disclosure, the resource determination sub-module includes a retrieval unit and a generation unit.

[0098] The retrieval unit is configured to retrieve the resource library according to the visual objects to obtain a target resource corresponding to the target content.

[0099] The generation unit is configured to call a large model to generate a response message including target resource information based on the visual objects.

[0100] According to an embodiment of the present disclosure, the content determination module 430 includes an intention determination sub-module and a content acquisition sub-module.

[0101] The intention determination sub-module is configured to determine the sub-intention of the target object, where the sub-intention includes at least one of obtaining the subsequent content of the target content in the target resource, obtaining the extended content of the target content, and obtaining the interpretation content of the target content.

[0102] The content acquisition sub-module is used to generate continuation content based on the target resource in response to the sub-intent, where the continuation content includes at least one of the subsequent content, extended content, and interpretation content of the target content.

[0103] According to an embodiment of the present disclosure, the intent determination sub-module includes a first resource determination unit and a first acquisition unit.

[0104] The first resource determination unit is used to determine the progress of the target content in the target resource according to the target image in response to the sub-intent being to obtain the subsequent content of the target content in the target resource.

[0105] The first acquisition unit is used to obtain the subsequent content of the target content in the target resource according to the progress.

[0106] According to an embodiment of the present disclosure, the intent determination sub-module includes a second resource determination unit and a second acquisition unit.

[0107] The second resource determination unit is used to determine the resource associated with the target resource in response to the sub-intent being to obtain the extended content of the target content.

[0108] The second acquisition unit is used to generate the extended content of the target content according to the target resource and the resource associated with the target resource.

[0109] According to an embodiment of the present disclosure, the intent determination sub-module includes a third acquisition unit, which is used to generate at least one of the summary and interpretation information of the target content according to the context of the target content in the target resource in response to the sub-intent being to obtain the interpretation content of the target content.

[0110] According to an embodiment of the present disclosure, the target content is at least one of book content, news content, and video content; the intent determination sub-module includes a fourth acquisition unit, which is used to obtain the audio of the subsequent content of the video content in the target resource as the continuation content according to the sub-intent in response to the target content being video content, or generate the summary information of the video content as the continuation content.

[0111] According to an embodiment of the present disclosure, the output module 440 includes a conversion sub-module and an output sub-module. The conversion sub-module is used to convert the continuation content into speech. The output sub-module is used to output the speech.

[0112] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0113] Figure 5FIG. 0 shows a schematic block diagram of an exemplary electronic device 500 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0114] As Figure 5 shown, the device 500 includes a computing unit 501 that may perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 may also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0115] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0116] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the information interaction method. For example, in some embodiments, the information interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the information interaction method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the information interaction method by any other suitable means (e.g., by means of firmware).

[0117] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0120] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0121] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0122] A computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0123] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0124] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An information interaction method, comprising: In response to determining that the intention of the target object is to obtain the continuation content of the target content according to the input information of the target object, using a wearable device to obtain a target image containing the target content; Determining a target resource corresponding to the target content according to the target image; Determining the continuation content of the target content according to the target resource; And Sending the continuation content to the wearable device so that the continuation content is output via the wearable device.

2. The method according to claim 1, wherein, The determining a target resource corresponding to the target content according to the target image includes: Extracting a visual object from the target image, where the visual object includes at least one of text, illustration, and video element; and Determining a target resource corresponding to the target content according to the visual object.

3. The method according to claim 2, wherein, The determining a target resource corresponding to the target content according to the visual object includes at least one of the following operations: Retrieving a resource library according to the visual object to obtain a target resource corresponding to the target content; Invoking a large model to generate a response message containing target resource information based on the visual object.

4. The method according to claim 1, wherein The determining the continuation content of the target content according to the target resource includes: Determining the sub-intention of the target object, where the sub-intention includes at least one of obtaining the subsequent content of the target content in the target resource, obtaining the extended content of the target content, and obtaining the interpretation content of the target content; In response to the sub-intention, generating the continuation content based on the target resource, where the continuation content includes at least one of the subsequent content, extended content, and interpretation content of the target content.

5. The method according to claim 4, wherein, The generating the continuation content based on the target resource in response to the sub-intention includes: In response to the sub-intention being to obtain the subsequent content of the target content in the target resource, determining the progress of the target content in the target resource according to the target image; and Obtaining the subsequent content of the target content in the target resource according to the progress.

6. The method according to claim 4, wherein The generating the continuation content based on the target resource in response to the sub-intention includes: In response to the sub-intention being to obtain the extended content of the target content, determining a resource associated with the target resource; and Generating the extended content of the target content according to the target resource and the resource associated with the target resource.

7. The method according to claim 4, wherein The generating the continuation content based on the target resource in response to the sub-intention includes: In response to the sub-intention being to obtain the interpretation content of the target content, generating at least one of a summary and interpretation information of the target content according to the context of the target content in the target resource.

8. The method according to claim 4, wherein The target content is at least one of book content, news content, and video content; the generating the continuation content based on the target resource in response to the sub-intention includes: In response to the target content being video content, obtaining the audio of the subsequent content of the video content in the target resource as the continuation content, or generating summary information of the video content as the continuation content.

9. The method according to claim 1, wherein Sending the continuation content to the wearable device for outputting the continuation content via the wearable device includes: Converting the continuation content into speech; and Outputting the speech via the wearable device.

10. An information interaction device, comprising: An image acquisition module, configured to, in response to determining, according to input information of a target object, that the intention of the target object is to obtain continuation content of a target content, obtain a target image including the target content by using a wearable device; A resource determination module, configured to determine a target resource corresponding to the target content according to the target image; A content determination module, configured to determine the continuation content of the target content according to the target resource; And An output module, configured to send the continuation content to the wearable device for outputting the continuation content via the wearable device.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, where the computer program is stored on at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method according to any one of claims 1 to 9.