Robot question and answer response generation method and device, computer equipment, readable storage medium and program product

By caching the scene description information in the robot question and answer system, the problem of low query processing efficiency in the prior art is solved, and the effect of reducing data processing volume and improving processing efficiency is achieved.

CN120014417AActive Publication Date: 2025-05-16SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510473575.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Every time the existing robot question and answer system enters a new query, it needs to decode the information in the scene corresponding to the big model, resulting in reduced efficiency.

Method used

By obtaining the scene description information of the current scene and cache it, the repeated calculation of the scene description information by the model when processing the query is reduced. The specific steps include obtaining the target scene image, processing the attention mechanism through the pre-trained model, obtaining the scene description information and cacheing it, and then directly using the cached scene description information when processing the query.

Benefits of technology

By caching scene description information, the data processing amount of the model during the query processing is reduced, the processing efficiency is improved, and the number of repeated calculations is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014417A_ABST
    Figure CN120014417A_ABST
Patent Text Reader

Abstract

The invention relates to a robot question and answer response generation method and device, computer equipment, a readable storage medium and a program product. The method comprises the following steps: acquiring scene description information corresponding to a current scene, wherein the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing based on a target scene image; receiving inquiry information; and inputting the inquiry information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the inquiry information, including: determining an attention result between the inquiry information and each element in a target scene image through the pre-trained model, and obtaining a response result corresponding to the inquiry information based on an attention processing result between any two elements and an attention result between the inquiry information and each element in the target scene image. According to the method, the scene description information is fully utilized, the data processing amount of the model is reduced, and the processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control technology, and in particular to a robot question and answer response generation method, device, computer equipment, readable storage medium and program product. Background Art

[0002] With the development of artificial intelligence technology, the application of robots in different scenarios has increased, especially in service, education, medical and other scenarios. Every time a user enters a new query, the scene and query are decoded and processed through large models such as GPT, and then the answer is obtained and output based on the processing result. In this way, the information in the scene corresponding to the large model needs to be decoded each time, which results in repeated operations and reduced efficiency. Summary of the invention

[0003] Based on this, it is necessary to provide a robot question and answer response generation method, device, computer equipment, computer-readable storage medium and computer program product that can reduce the data processing amount of the model and improve processing efficiency in response to the above technical problems.

[0004] In a first aspect, the present application provides a method for generating a robot question and answer response, the method comprising:

[0005] Obtain scene description information corresponding to the current scene, where the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on the target scene image;

[0006] Receive enquiry information;

[0007] The query information and the cached scene description information are input into a pre-trained model to obtain a response result corresponding to the query information.

[0008] In one embodiment, the obtaining scene description information corresponding to the current scene includes:

[0009] Acquire a target scene image;

[0010] Inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as scene description information;

[0011] The scene description information is cached.

[0012] In one embodiment, acquiring the target scene image includes:

[0013] Determine the difference between the current scene image and the historical scene image;

[0014] When the difference is greater than a threshold, the current scene image is used as a target scene image.

[0015] In one embodiment, the step of inputting the target scene image into a pre-trained model and using the processing result of the attention mechanism in the pre-trained model as the scene description information includes:

[0016] Extracting the scene sequence corresponding to the scene image through the feature extraction network of the pre-trained model;

[0017] Taking the scene sequence as the input of the decoder, and using the attention mechanism of the decoder to calculate the attention results between different elements in the scene sequence;

[0018] The attention result is used as scene description information.

[0019] In one embodiment, caching the scene description information includes:

[0020] The scene description information is stored in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is the attention result between any two elements.

[0021] In one embodiment, the step of inputting the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information includes:

[0022] When the response result is that no answer corresponding to the inquiry information is obtained, extracting keywords from the inquiry information;

[0023] Performing image segmentation on the target scene image based on the extracted keywords to obtain a region of interest;

[0024] The query information and the region of interest are input into a pre-trained model for processing to obtain a target response result.

[0025] In a second aspect, the present application further provides a robot question and answer response generating device, the device comprising:

[0026] A scene description information acquisition module, used to acquire scene description information corresponding to the current scene, wherein the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on the target scene image;

[0027] A receiving module, used for receiving inquiry information;

[0028] The response module is used to input the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information.

[0029] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in any one of the above-mentioned embodiments when executing the computer program.

[0030] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method in any one of the above-mentioned embodiments.

[0031] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps of the method in any one of the above embodiments when executed by a processor.

[0032] The above-mentioned robot question and answer response generation method, device, computer equipment, readable storage medium and program product first obtain the scene image, and obtain the scene description information based on the scene image, and cache it in advance. When making inquiries later, they are directly processed based on the query information and the scene description information without having to calculate the attention result between the scene description information again. This makes full use of the scene description information, reduces the data processing amount of the model, and improves processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0034] Figure 1 Schematic diagram of a flow chart of a method for generating a robot question and answer response in one embodiment;

[0035] Figure 2 is a flowchart of a step of acquiring scene description information in an embodiment;

[0036] Figure 3 is a processing flow chart of an embodiment in which a response result is that no answer corresponding to the inquiry information is obtained;

[0037] Figure 4 is a structural block diagram of a robot question and answer response generating device in one embodiment;

[0038] Figure 5FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0040] In one embodiment, Figure 1 As shown, a method for generating a robot question and answer response is provided. This embodiment uses the method applied to a robot terminal as an example. It can be understood that the method can also be applied to a server corresponding to the robot terminal, and can also be applied to a system including a robot terminal and a server, and is implemented through the interaction between the robot terminal and the server. In this embodiment, the method includes the following steps:

[0041] S102: Obtain scene description information corresponding to the current scene, where the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on the target scene image.

[0042] The scene description information is obtained by performing attention mechanism processing based on the target scene image, and the scene description information is used to describe the attention processing result between any two element tokens in the scene. The target scene image is an image of the scene in which the robot is located, and the robot performs image acquisition in real time during operation. In some optional embodiments, the robot may include multiple cameras, and each camera can capture a panoramic image of the scene in which the robot is located. In other embodiments, the robot may include only one camera or a binocular camera, and the field of view of the camera is the field of view of the robot, so that the image captured by the camera can represent the image of the scene in which the robot is located.

[0043] The attention mechanism is the attention mechanism of the decoder in the pre-trained model VLM, which is used to calculate the attention processing result between any two element tokens output by the feature extraction network.

[0044] S104: receiving inquiry information.

[0045] The inquiry information is information input by the user, for example, the inquiry information may be “what is included in the scene”, “what are the details of an object in the scene”, etc. The user may make inquiries based on needs.

[0046] After receiving the query information, the query information is processed to obtain a text sequence, for example, the text is segmented to obtain a vector corresponding to each segmented word, and then the vectors are connected to obtain a text sequence corresponding to the query information. In other embodiments, other methods can also be used for processing.

[0047] In addition, it should be noted that if the inquiry information is in the form of voice, in this application, the inquiry information is first converted from voice to text to obtain text information, and then the text information is processed.

[0048] S106: Input the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through a pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between the any two elements and the attention result between the query information and each element in the target scene image.

[0049] The pre-trained model VLM includes an attention mechanism, wherein the attention mechanism needs to calculate the attention results between any two element tokens. In this application, it is assumed that the scene description information includes the attention results between three element tokens, namely A, B and C. The scene description information includes the attention results between AB, the attention results between AC and the attention results between BC. Now there is query information. If the corresponding image-text conversion results are to be obtained, it is also necessary to calculate the attention results between the query information Q and the above three element tokens, that is, the attention results between AQ, the attention results between BQ and the attention results between CQ.

[0050] In the traditional technology, each time a query is received, the scene image and the query need to be input into the VLM, and the features of the scene image and the text of the query need to be extracted first. Then, the attention mechanism is calculated through transform. Taking the above text as an example, it is necessary to calculate the attention results between AB, the attention results between AC, the attention results between AQ, the attention results between BC, the attention results between BQ, and the attention results between CQ. These attention results need to be calculated each time a query is input, resulting in a large amount of repeated calculations.

[0051] In order to solve the above technical problems, the attention results between AB, the attention results between AC and the attention results between BC are stored as scene description information in this application. In this way, in subsequent processing, it is only necessary to calculate the attention results between each element token of the scene image and the query information Q. In this application, it is also necessary to cache each element token of the scene image. This can reduce the repeated calculation of the scene description information involved in the query processing summary and improve the processing efficiency.

[0052] The above-mentioned robot question and answer response generation method first obtains the scene image, and obtains the scene description information based on the scene image, and caches it in advance. When making inquiries later, it is directly processed based on the query information and the scene description information without having to calculate the attention result between the scene description information again. This makes full use of the scene description information, reduces the data processing amount of the model, and improves processing efficiency.

[0053] In one of the optional embodiments, in combination Figure 2 As shown, Figure 2 The present invention is a flowchart of a scene description information acquisition step in an embodiment. The scene description information acquisition step, i.e., the step of acquiring scene description information corresponding to the current scene, includes: acquiring a target scene image; inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as the scene description information; and caching the scene description information.

[0054] The pre-trained model VLM is a model that includes robot vision and a large prediction model, wherein the structure of the large prediction model can be a transform structure, so that the attention mechanism of the transform can be utilized, such as the cross-attention mechanism to calculate the cross-attention results of each element in the scene.

[0055] The pre-trained model VLM includes a feature extraction network and a decoder corresponding to a transform. The feature extraction network is used to extract the scene sequence corresponding to the scene image, and then the scene sequence is input into the transform for decoding. The decoding process of the transform includes the processing of the cross-attention mechanism, which can calculate the attention results between different elements in the scene sequence and cache the attention results as scene description information.

[0056] In the above embodiment, scene description information is generated in each new scene, so that when the query is subsequently processed, there is no need to process the scene image, but the scene description information and the query are directly processed, which reduces the repeated calculation of the scene description information and improves the calculation efficiency.

[0057] In one of the optional embodiments, the acquiring the target scene image includes: determining a difference between a current scene image and a historical scene image; and when the difference is greater than a threshold, using the current scene image as the target scene image.

[0058] The current scene image is a scene image captured by the robot's camera in real time, and the historical scene image is a scene image previously captured by the robot's camera. The historical scene image may be a target scene image of a previous scene, or any scene image captured from the previous scene to the current time. Optionally, the present application uses the historical scene image as an example of a target scene image of a previous scene for illustration.

[0059] Calculating the difference between the current scene image and the historical scene image may be based on slam. Specifically, calculating the difference between the current scene image and the historical scene image may be calculating the difference in feature points between the current scene image and the historical scene image. For example, the same portion in the current scene image and the historical scene image may be calculated, and then the ratio of the number of pixels in the same portion to the number of pixels in the current scene image may be used as the difference. In other embodiments, the difference may also be determined based on optical flow information, fixed objects, etc., which are not specifically limited here.

[0060] The threshold is preset, and the threshold may be obtained based on experience, such as 75%, etc. In other embodiments, other values ​​may also be selected, and no specific limitation is made here.

[0061] When the difference is greater than the threshold, it is determined that the scene has changed, and the target scene image corresponding to the scene needs to be obtained. Optionally, in order to facilitate processing in this application, the current scene image with a difference greater than the threshold is directly used as the target scene image, that is, the first frame scene image in the new scene.

[0062] In the above embodiment, the scene change can be judged based on the context scene image, thereby ensuring the real-time performance of the target scene image and further ensuring the accuracy of the scene description information.

[0063] In one of the optional embodiments, the target scene image is input into a pre-trained model, and the processing result of the attention mechanism in the pre-trained model is used as the scene description information, including: extracting a scene sequence corresponding to the scene image through a feature extraction network of the pre-trained model; using the scene sequence as the input of the decoder, and using the attention mechanism of the decoder to calculate the attention result between different elements in the scene sequence; and using the attention result as the scene description information.

[0064] Among them, the feature extraction network can be a simple neural network, which can extract the scene sequence corresponding to the scene image. Optionally, the feature extraction network can be connected to a fully connected layer. After the scene sequence is extracted by the scene extraction network, it is then classified through the fully connected layer to complete the prediction of each element token, and then the obtained element token is input into the decoder as a scene sequence.

[0065] The decoder is transform, and the decoder may include an attention mechanism. The attention mechanism may include a cross-attention mechanism, and the attention result between any two element tokens may be calculated through the cross-attention mechanism.

[0066] In the above embodiment, after entering a new scene, the scene description information is first calculated as the basis for subsequent inquiries. In this way, there is no need to calculate the attention results between any two element tokens in the scene in subsequent inquiries, which can improve processing efficiency.

[0067] In one of the optional embodiments, caching the scene description information includes: storing the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is the attention result between any two elements.

[0068] The scene description information includes the attention results between different element tokens, where the index can be the element token and the value is the attention result between any two element tokens.

[0069] In one of the optional embodiments, in combination Figure 3 As shown, Figure 3 The present invention is a processing flow chart for a case where a response result is that an answer corresponding to the query information is not obtained in an embodiment, wherein the query information and the cached scene description information are input into a pre-trained model to obtain a response result corresponding to the query information, including: when the response result is that an answer corresponding to the query information is not obtained, performing keyword extraction on the query information; performing image segmentation on the target scene image based on the extracted keywords to obtain a region of interest; and inputting the query information and the region of interest into a pre-trained model for processing to obtain a target response result.

[0070] Among them, the response result may include the answer corresponding to the inquiry information, and the answer that cannot be obtained corresponding to the inquiry information. If the response result is the answer corresponding to the inquiry information, the response result can be directly output. If the response result is that the answer corresponding to the inquiry information cannot be obtained, one method can output the answer that cannot be obtained. Another method is to generate response information in combination with context information when the response result is that the answer corresponding to the inquiry information cannot be obtained.

[0071] The keyword extraction of the query information can be performed through a large prediction model. After the target scene image is acquired, the target scene image is processed through the VLM, and the scene elements, that is, the element information corresponding to the element token, such as table, chair, etc., can be output.

[0072] By using the big prediction model to extract keywords from query information, keywords can be extracted from the query based on the scene elements output by the VLM, such as the keyword "table".

[0073] The target scene image is segmented based on the extracted keywords, which may be inputting the keywords and the target scene image into a segmentation model, which may be a SAM model or the like, which is not specifically limited here, and the target scene image is segmented based on the keywords by the segmentation model to determine the region of interest corresponding to the query information, and then the region of interest is segmented from the target scene image. This can reduce the amount of image processing.

[0074] Finally, the region of interest and query information are input into the VLM for processing, where the feature extraction network of the VLM is used to extract the image sequence of the region of interest, and the text processing module is used to extract the text features of the query information to obtain the text sequence. The image sequence and text sequence are then input into the VLM decoder, and the attention mechanism of the VLM decoder is used to perform attention processing on the image sequence and text sequence, and finally the target response result is obtained.

[0075] In the above embodiment, when no answer corresponding to the query information is obtained, the contextual query information can be used to segment the scene image to obtain the region of interest, and then VLM processing is performed based on the query information and the scene image to obtain the target response result, thereby making full use of the contextual information to make the robot's response result more accurate.

[0076] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0077] Based on the same inventive concept, the embodiment of the present application also provides a robot question and answer response generation device for implementing the robot question and answer response generation method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more robot question and answer response generation device embodiments provided below can refer to the limitations of the robot question and answer response generation method above, and will not be repeated here.

[0078] In an exemplary embodiment, Figure 4 As shown, a robot question and answer response generating device is provided, comprising: a scene description information acquisition module 401, a receiving module 402 and a response module 403, wherein:

[0079] A scene description information acquisition module 401 is used to acquire scene description information corresponding to the current scene, where the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on the target scene image;

[0080] Receiving module 402, used to receive inquiry information;

[0081] The response module 403 is used to input the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through a pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between any two elements and the attention result between the query information and each element in the target scene image.

[0082] In one of the optional embodiments, the scene description information acquisition module 401 is specifically used to acquire a target scene image; input the target scene image into a pre-trained model, and use the processing result of the attention mechanism in the pre-trained model as the scene description information; and cache the scene description information.

[0083] In one of the optional embodiments, the scene description information acquisition module 401 is specifically used to determine the difference between the current scene image and the historical scene image; when the difference is greater than a threshold, the current scene image is used as the target scene image.

[0084] In one of the optional embodiments, the scene description information acquisition module 401 is specifically used to extract a scene sequence corresponding to a scene image through a feature extraction network of a pre-trained model; use the scene sequence as the input of a decoder, and use the attention mechanism of the decoder to calculate the attention results between different elements in the scene sequence; and use the attention results as scene description information.

[0085] In one of the optional embodiments, the scene description information acquisition module 401 is specifically used to store the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is the attention result between any two elements.

[0086] In one of the optional embodiments, the above-mentioned response module 403 is specifically used to extract keywords from the query information when the response result is that no answer corresponding to the query information is obtained; perform image segmentation on the target scene image based on the extracted keywords to obtain a region of interest; and input the query information and the region of interest into a pre-trained model for processing to obtain a target response result.

[0087] Each module in the above-mentioned robot question and answer response generating device can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0088] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a robot question and answer response generation method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0089] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0090] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the following steps when executing the computer program: obtaining scene description information corresponding to a current scene, wherein the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on a target scene image; receiving query information; inputting the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through a pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between the any two elements and the attention result between the query information and each element in the target scene image.

[0091] In one embodiment, the step of obtaining scene description information corresponding to the current scene implemented when the processor executes a computer program includes: obtaining a target scene image; inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as the scene description information; and caching the scene description information.

[0092] In one embodiment, the acquiring of the target scene image implemented when the processor executes a computer program includes: determining the difference between the current scene image and the historical scene image; and when the difference is greater than a threshold, using the current scene image as the target scene image.

[0093] In one embodiment, the step of inputting the target scene image into a pre-trained model and using the processing result of the attention mechanism in the pre-trained model as scene description information implemented when the processor executes a computer program includes: extracting a scene sequence corresponding to the scene image through a feature extraction network of the pre-trained model; using the scene sequence as input to a decoder, and using the attention mechanism of the decoder to calculate the attention result between different elements in the scene sequence; and using the attention result as the scene description information.

[0094] In one embodiment, the caching of the scene description information implemented when the processor executes a computer program includes: storing the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is an attention result between any two elements.

[0095] In one embodiment, the step of inputting the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information implemented when a processor executes a computer program includes: when the response result is that no answer corresponding to the query information is obtained, performing keyword extraction on the query information; performing image segmentation on the target scene image based on the extracted keywords to obtain a region of interest; and inputting the query information and the region of interest into a pre-trained model for processing to obtain a target response result.

[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program implements the following steps when executed by a processor: obtaining scene description information corresponding to a current scene, the scene description information being an attention processing result between any two elements obtained by performing attention mechanism processing on a target scene image; receiving query information; inputting the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through a pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between the any two elements and the attention result between the query information and each element in the target scene image.

[0097] In one embodiment, the method of obtaining scene description information corresponding to the current scene implemented when a computer program is executed by a processor includes: obtaining a target scene image; inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as the scene description information; and caching the scene description information.

[0098] In one embodiment, the acquisition of the target scene image implemented when the computer program is executed by the processor includes: determining the difference between the current scene image and the historical scene image; and when the difference is greater than a threshold, using the current scene image as the target scene image.

[0099] In one embodiment, the method implemented when a computer program is executed by a processor includes: inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as scene description information, including: extracting a scene sequence corresponding to the scene image through a feature extraction network of the pre-trained model; using the scene sequence as input to a decoder, and using the attention mechanism of the decoder to calculate the attention result between different elements in the scene sequence; and using the attention result as the scene description information.

[0100] In one embodiment, the caching of the scene description information implemented when the computer program is executed by a processor includes: storing the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is an attention result between any two elements.

[0101] In one embodiment, the computer program implemented when executed by a processor inputs the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: when the response result is that no answer corresponding to the query information is obtained, performing keyword extraction on the query information; performing image segmentation on the target scene image based on the extracted keywords to obtain a region of interest; and inputting the query information and the region of interest into a pre-trained model for processing to obtain a target response result.

[0102] In one embodiment, a computer program product is provided, including a computer program, which implements the following steps when executed by a processor: obtaining scene description information corresponding to a current scene, the scene description information being an attention processing result between any two elements obtained by performing attention mechanism processing on a target scene image; receiving query information; inputting the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through the pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between the any two elements and the attention result between the query information and each element in the target scene image.

[0103] In one embodiment, the method of obtaining scene description information corresponding to the current scene implemented when a computer program is executed by a processor includes: obtaining a target scene image; inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as the scene description information; and caching the scene description information.

[0104] In one embodiment, the acquisition of the target scene image implemented when the computer program is executed by the processor includes: determining the difference between the current scene image and the historical scene image; and when the difference is greater than a threshold, using the current scene image as the target scene image.

[0105] In one embodiment, the method implemented when a computer program is executed by a processor includes: inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as scene description information, including: extracting a scene sequence corresponding to the scene image through a feature extraction network of the pre-trained model; using the scene sequence as input to a decoder, and using the attention mechanism of the decoder to calculate the attention result between different elements in the scene sequence; and using the attention result as the scene description information.

[0106] In one embodiment, the caching of the scene description information implemented when the computer program is executed by a processor includes: storing the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is an attention result between any two elements.

[0107] In one embodiment, the computer program implemented when executed by a processor inputs the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: when the response result is that no answer corresponding to the query information is obtained, performing keyword extraction on the query information; performing image segmentation on the target scene image based on the extracted keywords to obtain a region of interest; and inputting the query information and the region of interest into a pre-trained model for processing to obtain a target response result.

[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0109] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0110] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0111] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A robot question and answer response generation method, characterized in that: The method comprises: Obtain scene description information corresponding to the current scene, where the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on the target scene image; Receive enquiry information; The query information and the cached scene description information are input into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through a pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between the any two elements and the attention result between the query information and each element in the target scene image.

2. The method according to claim 1, characterized in that: The obtaining of scene description information corresponding to the current scene includes: Acquire a target scene image; Inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as scene description information; The scene description information is cached.

3. The method according to claim 2, characterized in that The step of acquiring the target scene image comprises: Determine the difference between the current scene image and the historical scene image; When the difference is greater than a threshold, the current scene image is used as a target scene image.

4. The method according to claim 2, characterized in that: The step of inputting the target scene image into a pre-trained model and using the processing result of the attention mechanism in the pre-trained model as scene description information includes: Extracting the scene sequence corresponding to the scene image through the feature extraction network of the pre-trained model; Taking the scene sequence as the input of the decoder, and using the attention mechanism of the decoder to calculate the attention results between different elements in the scene sequence; The attention result is used as scene description information.

5. The method according to claim 2, characterized in that: The caching the scene description information includes: The scene description information is stored in a k-value cache, wherein the index k of the scene description information in the k-value cache is an element, and the value value is the attention result between any two elements.

6. The method according to any one of claims 1 to 5, characterized in that: The step of inputting the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information includes: When the response result is that no answer corresponding to the inquiry information is obtained, extracting keywords from the inquiry information; Performing image segmentation on the target scene image based on the extracted keywords to obtain a region of interest; The query information and the region of interest are input into a pre-trained model for processing to obtain a target response result.

7. A robot question and answer response generating device, characterized in that: The device comprises: A scene description information acquisition module, used to acquire scene description information corresponding to the current scene, wherein the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing on the target scene image; A receiving module, used for receiving inquiry information; A response module is used to input the query information and the cached scene description information into a pre-trained model to obtain a response result corresponding to the query information, including: determining the attention result between the query information and each element in the target scene image through a pre-trained model, and obtaining a response result corresponding to the query information based on the attention processing result between the any two elements and the attention result between the query information and each element in the target scene image.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Image-based question and answer method and device, equipment, storage medium and program product

    CN117312508A

  • Robot control system and method, storage medium, controller and robot

    CN118927246A

  • Multi-modal question and answer method and device, electronic equipment and computer readable storage medium

    CN119047583A

  • Method and system for automated visual question answering

    EP3920048A1

  • Unified Vision and Dialogue Transformer with BERT

    US20210232773A1