Interaction processing method and device, electronic equipment, readable storage medium and program product

By displaying images captured by the camera and collecting multimodal information in the interface, and using a smart assistant application to analyze the images and user input information, the problem of low efficiency in traditional interactive processing is solved, and the effect of quickly obtaining analysis results is achieved.

CN121597089APending Publication Date: 2026-03-03GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411156060.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Traditional interactive processing methods are inefficient and cannot quickly acquire and analyze multimodal information.

Method used

By displaying images captured by the camera in the interface and including information collection controls, the system responds to user operations to collect multimodal information and uses a smart assistant application to analyze the images and user input information to obtain analysis results.

Benefits of technology

It enables the rapid acquisition of analysis results from multimodal information, improving the efficiency and accuracy of interactive processing, especially in scenarios involving rapid movement or requiring quick responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597089A_ABST
    Figure CN121597089A_ABST
Patent Text Reader

Abstract

The invention relates to an interaction processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps of displaying a first interface in response to a trigger operation, and displaying a shot image acquired through a camera in the first interface; the first interface further comprises an information acquisition control; in response to a trigger operation on the information acquisition control, acquiring information input by a user; and obtaining an analysis result according to the shot image and the information input by the user, and presenting the analysis result. By adopting the method, the interactive processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software interaction technology, and in particular to an interaction processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] With the development of artificial intelligence (AI) technology, more and more AI applications are emerging, constantly expanding the boundaries of human-computer interaction and making interactions more intelligent, personalized, and efficient. As technology continues to advance, the application of AI in human-computer interaction will become more widespread and in-depth in the future, significantly changing people's lifestyles and work patterns. Users can interact with AI applications, making life more convenient and work more efficient.

[0003] However, traditional interactive processing methods suffer from low efficiency. Summary of the Invention

[0004] This application provides an interactive processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the efficiency of interactive processing.

[0005] Firstly, this application provides an interactive processing method, including:

[0006] In response to a trigger operation, a first interface is displayed, and the captured image obtained through the camera is displayed on the first interface; the first interface also includes an information collection control;

[0007] In response to a trigger operation on the information collection control, information input by the user is collected;

[0008] Based on the captured images and the information input by the user, the analysis results are obtained and presented.

[0009] In one embodiment, the method further includes:

[0010] In response to a trigger operation of the information acquisition control, a target image is acquired from each of the captured images;

[0011] The step of obtaining analysis results based on the captured image and the information input by the user includes:

[0012] The analysis results are obtained based on the target image and the information input by the user.

[0013] In one embodiment, the step of acquiring a target image from each of the captured images in response to a trigger operation of the information acquisition control includes:

[0014] In response to the triggering operation of the information acquisition control, the captured image corresponding to the triggering of the information acquisition control is determined from each of the captured images, and the target image is obtained based on the captured image at the triggering of the information acquisition control.

[0015] In one embodiment, the method further includes:

[0016] During the process of collecting user input information, the target image is displayed on the first interface.

[0017] In one embodiment, the information acquisition control includes a voice acquisition control; the step of acquiring user-input information in response to a trigger operation on the information acquisition control includes:

[0018] In response to a trigger operation on the voice acquisition control, the target voice input by the user is acquired via the microphone.

[0019] In one embodiment, the step of acquiring the target speech input by the user via a microphone in response to a trigger operation on the speech acquisition control includes:

[0020] In response to a trigger operation on the voice acquisition control, candidate signals input by the user are acquired via the microphone;

[0021] The candidate signal is subjected to speech activity detection. If the candidate signal is determined to be a speech signal, then the candidate signal is used as the target speech input by the user.

[0022] In one embodiment, after displaying the captured image via the camera on the first interface, the method further includes:

[0023] Optical character recognition is performed on the captured image to obtain the captured image text;

[0024] The step of obtaining analysis results based on the captured image and the information input by the user includes:

[0025] The analysis results are obtained based on the captured image, the captured image text, and the information input by the user.

[0026] Secondly, this application also provides an interactive processing apparatus, comprising:

[0027] The first interface display module is used to respond to a trigger operation, display a first interface, and display the captured image obtained by the camera on the first interface; the first interface also includes an information collection control;

[0028] The user input information collection module is used to collect user input information in response to the trigger operation of the information collection control;

[0029] The analysis result acquisition module is used to acquire analysis results based on the captured image and the information input by the user, and to present the analysis results.

[0030] Thirdly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0031] In response to a trigger operation, a first interface is displayed, and the captured image obtained through the camera is displayed on the first interface; the first interface also includes an information collection control;

[0032] In response to a trigger operation on the information collection control, information input by the user is collected;

[0033] Based on the captured images and the information input by the user, the analysis results are obtained and presented.

[0034] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0035] In response to a trigger operation, a first interface is displayed, and the captured image obtained through the camera is displayed on the first interface; the first interface also includes an information collection control;

[0036] In response to a trigger operation on the information collection control, information input by the user is collected;

[0037] Based on the captured images and the information input by the user, the analysis results are obtained and presented.

[0038] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0039] In response to a trigger operation, a first interface is displayed, and the captured image obtained through the camera is displayed on the first interface; the first interface also includes an information collection control;

[0040] In response to a trigger operation on the information collection control, information input by the user is collected;

[0041] Based on the captured images and the information input by the user, the analysis results are obtained and presented.

[0042] The aforementioned interactive processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, in response to a trigger operation, display a first interface and show an image captured by a camera on the first interface. The first interface also includes an information acquisition control. In response to a trigger operation on the information acquisition control, user-inputted information is collected. That is, multimodal information, such as the captured image and user-inputted information, can be quickly obtained on the first interface without having to collect information separately by calling a camera application and other information applications and then inserting it into the application. Therefore, based on the captured image and user-inputted information, analysis results can be obtained and presented more quickly, improving the efficiency of interactive processing. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart illustrating the interaction processing method in one embodiment;

[0045] Figure 2 This is a flowchart illustrating the interaction processing method in another embodiment;

[0046] Figure 3 This is a schematic diagram of a mobile phone desktop in one embodiment;

[0047] Figure 4 This is a schematic diagram of a smart assistant page including an information interaction entry point in one embodiment;

[0048] Figure 5 This is a schematic diagram of a smart assistant page including a multimodal input field in one embodiment;

[0049] Figure 6 This is a schematic diagram of the first interface in one embodiment;

[0050] Figure 7 This is a schematic diagram of the page when the voice acquisition control is triggered in one embodiment;

[0051] Figure 8 This is a schematic diagram of the analysis results page in one embodiment;

[0052] Figure 9 This is a structural block diagram of the interactive processing device in one embodiment;

[0053] Figure 10This is a diagram of the internal structure of an electronic device in one embodiment. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0055] In one embodiment, such as Figure 1 As shown, an interactive processing method is provided. This embodiment illustrates the application of this method to an electronic device, which may be a terminal or a server.

[0056] It is understandable that this interactive processing method can also be applied to systems including terminals and servers, and implemented through the interaction between the terminal and the server. For example, in response to a trigger operation, the terminal displays a first interface, which shows an image captured by a camera. This first interface also includes an information collection control. In response to a trigger operation on the information collection control, the terminal collects information input by the user, sends the captured image and the user input information to the server, and instructs the server to obtain analysis results based on the captured image and the user input information, and return the analysis results to the terminal. The terminal receives the analysis results returned by the server and presents the analysis results.

[0057] Electronic devices can include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can include virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing cloud computing services.

[0058] In this embodiment, applied to an electronic device, the interaction processing method includes the following steps:

[0059] In step S102, in response to the trigger operation, a first interface is displayed, and the captured image obtained by the camera is displayed on the first interface; the first interface also includes an information collection control.

[0060] The trigger operation is used to display the first interface. For example, the trigger operation can be a wake-up operation of the smart assistant application or a pressing operation of a specified physical button, and is not limited to these. The wake-up operation of the smart assistant application can be voice wake-up, gesture wake-up, button wake-up, or other wake-up operations, and is not limited here.

[0061] The first interface is used to collect multimodal information. It includes a shooting area for displaying captured images and information acquisition controls for collecting user input. The captured images can be a single frame or at least two frames.

[0062] Optionally, the electronic device displays a first interface in response to a triggering action on the smart assistant application.

[0063] The Smart Assistant application is an application that integrates artificial intelligence technology, aiming to provide users with convenient, efficient, and intelligent life services. It can provide video analysis, image analysis, and text analysis; it can also generate various videos, images, and text; and it can perform intelligent retrieval and other processing.

[0064] Optionally, the electronic device displays a multimodal input field in response to an activation operation of the smart assistant application; and displays a first interface in response to a trigger operation of the multimodal input field.

[0065] A multimodal input entry point is used to receive multimodal information input from the user. Multimodal information includes information of various different modalities, such as text, images, audio, and video.

[0066] Optionally, in response to a wake-up operation of the smart assistant application, a multimodal input entry is displayed, including: in response to a wake-up operation of the smart assistant application, a smart assistant page is displayed; the smart assistant page includes an information interaction entry; and in response to a trigger operation of the information interaction entry, a multimodal input entry is displayed.

[0067] The Smart Assistant page is the initial page displayed after launching the Smart Assistant application. It displays the information interaction entry point, as well as keyboard input and other information. The information interaction entry point is used to receive input information for interaction between the user and the Smart Assistant application.

[0068] Optionally, in response to a triggering operation on an information interaction entry point, the electronic device may display a multimodal input entry point, and may also display at least one of a question-and-answer entry point, a document analysis entry point, and an image parsing entry point.

[0069] Optionally, in response to a trigger operation on the multimodal input port, the electronic device displays a first interface, calls the camera module in the electronic device to acquire a captured image, and displays the captured image in the shooting area of ​​the first interface.

[0070] It is understandable that the first interface includes a shooting area and information collection controls. The position and size of the shooting area and information collection controls on the first interface can be set as needed, and there are no restrictions here.

[0071] Optionally, the electronic device can display the captured image in real time in the shooting area of ​​the first interface.

[0072] Optionally, the electronic device displays a single frame of the target image in the shooting area of ​​the first interface.

[0073] Step S104: In response to the trigger operation of the information collection control, collect the information input by the user.

[0074] The information input by the user can be voice, text, or other types of information, and there are no restrictions on this.

[0075] Optionally, the electronic device, in response to a triggering operation of the information acquisition control, acquires at least one of voice information and text information.

[0076] For example, the electronic device displays the captured image in the shooting area of ​​the first interface. In response to the triggering operation of the information collection control, the user can input information while viewing the captured image, and the electronic device collects the information input by the user.

[0077] Step S106: Based on the captured images and the information input by the user, obtain the analysis results and present the analysis results.

[0078] The analysis results are obtained by analyzing the captured images and the information input by the user.

[0079] Optionally, the electronic device sends the captured image and user-inputted information to a smart assistant application, which then analyzes the captured image and user-inputted information to obtain analysis results.

[0080] Optionally, the electronic device displays the analysis results in the shooting area of ​​the first interface.

[0081] Optionally, the electronic device presents the analysis results in a perceptible manner. Perceptible methods include text display and voice playback.

[0082] For example, after a user activates the smart assistant application, the electronic device displays a multimodal input field. In response to the user's triggering operation on the multimodal input field, a first interface is displayed, and the animal that the user wants to photograph is displayed in the shooting area of ​​the first interface. In response to the triggering operation of the voice acquisition control, the information collected from the user's input is the user's question "What animal is this?" Then, the smart assistant application will perform intent parsing based on the captured image and the user's input information, thereby obtaining the analysis results and presenting the analysis results.

[0083] The aforementioned interactive processing method, in response to a trigger operation, displays a first interface showing the image captured by the camera. The first interface also includes an information collection control. In response to a trigger operation on the information collection control, it collects user-inputted information. That is, multimodal information, including the captured image and user-inputted information, can be quickly obtained within this first interface, eliminating the need to separately collect information from camera applications and other information applications and then insert it into the application. Therefore, based on the captured image and user-inputted information, analysis results can be obtained and presented more quickly, improving the efficiency of interactive processing. Furthermore, in scenarios involving rapid movement or requiring quick camera response, the aforementioned interactive processing method can also quickly interact with and ask questions of the intelligent assistant, further improving the efficiency of interactive processing. It allows users to promptly ask questions about what they see, improving the efficiency of resolving user problems and enhancing the user experience, thereby increasing the convenience of interaction between the intelligent assistant application and the user's perspective.

[0084] In one embodiment, the method further includes: in response to a trigger operation of the information acquisition control, acquiring a target image from each captured image; and acquiring analysis results based on the captured images and user input information, including: acquiring analysis results based on the target image and user input information.

[0085] The target image is the image displayed in the shooting area after the information acquisition control is triggered.

[0086] Optionally, the electronic device responds to a trigger operation of the information acquisition control to acquire the target image from each captured image and collect information input by the user.

[0087] Optionally, the electronic device determines the timestamp that triggers the operation and obtains the target image corresponding to that timestamp from each captured image.

[0088] Optionally, the electronic device determines the timestamp that triggers the operation and the time range within which the timestamp falls; it acquires candidate images within that time range from each captured image; and it determines the target image from each candidate image. The electronic device may randomly determine the target image from each candidate image, or it may select the candidate image with the highest clarity from each candidate image as the target image, or it may use other methods to determine the target image, which are not limited here.

[0089] Optionally, the electronic device sends the target image and user-inputted information to a smart assistant application, which then analyzes the target image and user-inputted information to obtain analysis results.

[0090] Optionally, in response to a trigger operation of the information acquisition control, a target image is acquired from each captured image, including: in response to a trigger operation of the information acquisition control, determining the captured image corresponding to the trigger of the information acquisition control from each captured image, and acquiring the target image based on the captured image at the trigger of the information acquisition control.

[0091] Optionally, the electronic device displays each captured image in real time in the shooting area of ​​the first interface, and in response to the triggering operation of the information acquisition control, determines the captured image corresponding to the triggering of the information acquisition control from each captured image, and obtains the target image from the captured image at the triggering of the information acquisition control.

[0092] Optionally, the electronic device can perform downsampling, cropping, or other operations on the image captured when the information acquisition control is triggered to obtain the target image.

[0093] Optionally, the above method further includes: displaying the target image in the shooting area of ​​the first interface during the process of collecting user input information.

[0094] Understandably, in response to the triggering operation of the information acquisition control, the electronic device displays the target image in the shooting area of ​​the first interface, while the user inputs relevant descriptive information based on the target image. Thus, the electronic device can simultaneously acquire user input information that is more relevant to the target image, improving the accuracy of interactive processing.

[0095] In this embodiment, the electronic device responds to the trigger operation of the information acquisition control to acquire the target image from each captured image. It can analyze the captured screen composed of multiple images. That is, based on the target image and the information input by the user, the analysis results can be obtained more quickly, thus improving the efficiency of interactive processing.

[0096] In one embodiment, the information acquisition control includes a voice acquisition control; in response to a trigger operation on the information acquisition control, acquiring user input information includes: in response to a trigger operation on the voice acquisition control, acquiring target voice input by the user through a microphone.

[0097] The voice acquisition control is used to acquire the target voice.

[0098] Optionally, the information acquisition control includes at least one of a voice acquisition control and a text acquisition control. A text acquisition control is a control used to acquire target text.

[0099] Optionally, in response to a triggering operation on the text acquisition control, the electronic device displays a virtual keyboard and acquires the target text via the virtual keyboard. Both the target speech and the target text are information input by the user.

[0100] For example, the electronic device displays a first interface, a voice acquisition control on the first interface, and a shooting area in the first interface displays a real-time captured image by the user; in response to a trigger operation of the voice acquisition control, the target voice of the user describing the captured image is acquired through a microphone.

[0101] For example, the electronic device displays a first interface, a text acquisition control on the first interface, and a shooting area in the first interface displays a real-time captured image by the user; in response to a trigger operation on the text acquisition control, a virtual keyboard is displayed, and the target text described by the user for the captured image is acquired through the virtual keyboard.

[0102] Optionally, in response to a trigger operation of the voice acquisition control, the target voice input by the user is acquired via a microphone, including: in response to a trigger operation of the voice acquisition control, acquiring a candidate signal input by the user via a microphone; performing voice activity detection on the candidate signal; and if the candidate signal is determined to be a voice signal, then using the candidate signal as the target voice input by the user.

[0103] Voice activity detection (VAD), also known as speech activity detection or speech detection, is a technique used in speech processing to detect the presence of speech signals.

[0104] Optionally, the electronic device performs voice activity detection on the candidate signal. If the candidate signal is determined to be a voice signal, it is used as the target voice input by the user. If the candidate signal is determined to be a non-voice signal, it is discarded and the user continues to collect candidate signals input by the microphone.

[0105] In this embodiment, the electronic device, in response to a trigger operation of the voice acquisition control, acquires the target voice input by the user through a microphone. Based on the captured image and the user-input target voice, analysis results can be obtained more quickly. Furthermore, the electronic device, through voice activity detection, can accurately acquire the target voice signal, thereby more accurately analyzing the captured image and target voice to obtain more precise analysis results.

[0106] In one embodiment, after displaying the captured image through the camera on the first interface, the method further includes: performing optical character recognition on the captured image to obtain captured image text; and obtaining analysis results based on the captured image and user input information, including: obtaining analysis results based on the captured image, captured image text, and user input information.

[0107] Optical Character Recognition (OCR) refers to the process by which electronic devices determine the shape of characters by detecting dark and light patterns, and then translate the shapes into computer text using character recognition methods. Image text refers to text captured in an image.

[0108] Optionally, the electronic device sends the captured image, the captured image text, and the user-inputted information to a smart assistant application, which then analyzes the captured image, the captured image text, and the user-inputted information to obtain the analysis results.

[0109] Optionally, during the process of collecting user input information, the electronic device obtains a first analysis result based on the captured image and text; after collecting user input information, it obtains a second analysis result based on the captured image and user input information; and based on the first and second analysis results, it obtains an analysis result.

[0110] Optionally, the electronic device sends the captured image and text to a smart assistant application, which analyzes the captured image and text to obtain a first analysis result; after collecting information input by the user, it sends the captured image and the user input information to the smart assistant application, which analyzes the captured image and the user input information to obtain a second analysis result; the first analysis result and the second analysis result are combined to obtain the final analysis result.

[0111] It is understandable that the information input by the user may take a long time. Therefore, in the process of collecting the user's input information, the electronic device can first analyze the captured image and text to obtain the first analysis result. After collecting the user's input information, it can obtain the second analysis result based on the captured image and the user's input information, and then obtain the complete analysis result, which can improve the efficiency of the analysis.

[0112] For example, the electronic device displays a captured image of product A in the shooting area of ​​the first interface. Optical character recognition (OCR) is performed on the captured image to obtain text representing product A. In response to a trigger operation of the information collection control, the user input information is collected as a purchase link for product A. During the collection of user input information, the electronic device can first analyze the text of product A to obtain relevant information about product A. After collecting the user input information, a second analysis result is obtained based on the captured image and the user input information, resulting in a purchase link for product A. Therefore, based on the first and second analysis results, the final analysis result is the relevant information about product A and the purchase link for product A.

[0113] In this embodiment, the electronic device performs optical character recognition on the captured image to obtain the captured image text. Based on the captured image, the captured image text, and the information input by the user, the analysis can be performed more accurately to obtain more accurate analysis results.

[0114] In one embodiment, another interaction processing method is also provided, applied to an electronic device, the interaction processing method comprising the following steps:

[0115] Step S202: Launch the smart assistant application.

[0116] Step S204: Activate the multimodal input capability. That is, the electronic device displays the multimodal input interface.

[0117] Step S206: Click the multimodal input button.

[0118] Step S208: Activate the shooting screen.

[0119] Step S210: Click the voice ball to take a screenshot and start the conversation. The voice ball is the voice capture control.

[0120] Step S212: The user inputs the target voice.

[0121] In step S214, if the voice activity detection determines that it is a voice signal, then the target image and target voice are sent to the intelligent assistant application. The target image is obtained from the captured image.

[0122] Step S216: The smart assistant application generates the analysis results.

[0123] In one embodiment, a mobile phone is used as the electronic device for illustration, with reference to... Figures 3 to 8 , Figure 3This is a screenshot of the phone's home screen. When the home screen is displayed, in response to a request to activate the smart assistant application, the smart assistant page is displayed. This smart assistant page includes an information interaction entry point, as shown below. Figure 4 As shown. In response to a trigger operation on the information interaction entry point, the mobile phone displays a multimodal input entry point, including the intelligent assistant page with the multimodal input entry point, as shown below. Figure 5 As shown. In response to the trigger operation of the multimodal input field, the mobile phone displays the first interface, as shown below. Figure 6 As shown, the captured image is displayed in the shooting area of ​​the first interface, and a voice capture control is displayed at the bottom of the first interface. In response to the triggering operation of the voice capture control, the phone determines the captured image corresponding to the triggering of the voice capture control from among the captured images, obtains the target image based on the captured image at the time of triggering the voice capture control, and displays the target image in the shooting area. Additionally, it captures the target voice input by the user through the microphone. The page when the voice capture control is triggered is as follows: Figure 7 As shown in the image, the phone sends the target image and target voice to the smart assistant application and waits for the analysis results generated by the application. The smart assistant application generates an analysis based on the target image and target voice and displays the results, as shown in the image. Figure 8 As shown.

[0124] In one embodiment, another interaction processing method is also provided, applied to an electronic device, the interaction processing method comprising the following steps:

[0125] Step A1: In response to the trigger operation, a first interface is displayed, and the captured image is shown in the first interface through the camera; the first interface includes information collection controls.

[0126] Step A2: Perform optical character recognition on the captured image to obtain the text in the captured image.

[0127] Step A3: In response to the triggering operation of the voice acquisition control, determine the image corresponding to the triggering of the information acquisition control from each captured image, acquire the target image based on the captured image when the information acquisition control is triggered, and acquire the target voice input by the user through the microphone. During the acquisition of the target voice, the target image is displayed in the capture area of ​​the first interface.

[0128] The process of acquiring target speech via microphone includes: acquiring candidate signals via microphone; performing speech activity detection on the candidate signals; and if the candidate signals are determined to be speech signals, then using the candidate signals as target speech.

[0129] Step A4: Obtain the analysis results based on the target image, the captured image text, and the target speech.

[0130] Step A5 presents the analysis results.

[0131] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0132] Based on the same inventive concept, this application also provides an interactive processing apparatus for implementing the interactive processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more interactive processing apparatus embodiments provided below can be found in the limitations of the interactive processing method described above, and will not be repeated here.

[0133] In one exemplary embodiment, such as Figure 9 As shown, an interactive processing device is provided, including: a first interface display module 902, a user input information acquisition module 904, and an analysis result acquisition module 906, wherein:

[0134] The first interface display module 902 is used to display the first interface in response to the trigger operation, and to display the captured image obtained by the camera in the first interface; the first interface also includes an information collection control.

[0135] The user input information collection module 904 is used to collect user input information in response to the trigger operation of the information collection control.

[0136] The analysis result acquisition module 906 is used to acquire and present the analysis results based on the captured images and user input information.

[0137] The aforementioned interactive processing device, in response to a trigger operation, displays a first interface and shows the captured image obtained through the camera on the first interface. The first interface also includes an information collection control. In response to a trigger operation on the information collection control, it collects information input by the user. That is, multimodal information such as the captured image and the user input information can be quickly obtained on the first interface without having to call the camera application and other information applications to collect information and then insert it into the application. Therefore, based on the captured image and the user input information, the analysis results can be obtained and presented more quickly, improving the efficiency of interactive processing.

[0138] In one embodiment, the first interface display module 902 is further configured to acquire a target image from various captured images in response to a trigger operation of the information acquisition control; the analysis result acquisition module 906 is further configured to acquire analysis results based on the target image and the information input by the user.

[0139] In one embodiment, the first interface display module 902 is further configured to, in response to a trigger operation of the information acquisition control, determine the captured image corresponding to the trigger of the information acquisition control from each captured image, and acquire the target image based on the captured image at the trigger of the information acquisition control.

[0140] In one embodiment, the first interface display module 902 is further configured to display the target image in the shooting area of ​​the first interface during the process of collecting user input information.

[0141] In one embodiment, the information acquisition control includes a voice acquisition control; the user input information acquisition module 904 is further configured to acquire the user's target voice input via a microphone in response to a trigger operation on the voice acquisition control.

[0142] In one embodiment, the user input information acquisition module 904 is further configured to, in response to a trigger operation of the voice acquisition control, acquire candidate signals input by the user through a microphone; perform voice activity detection on the candidate signals; and if the candidate signals are determined to be voice signals, use the candidate signals as the target voice input by the user.

[0143] In one embodiment, the above-mentioned device further includes an optical character recognition module for performing optical character recognition on the captured image to obtain captured image text; the above-mentioned analysis result acquisition module 906 is also used to acquire analysis results based on the captured image, captured image text and user input information.

[0144] Each module in the aforementioned interactive processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to each module.

[0145] In one exemplary embodiment, an electronic device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, this electronic device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an interactive processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the electronic device, or external keyboards, touchpads, or mice, etc.

[0146] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0147] In one exemplary embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: in response to a trigger operation, displaying a first interface and displaying a captured image obtained by a camera on the first interface; the first interface further includes an information acquisition control; in response to a trigger operation on the information acquisition control, acquiring information input by the user; and based on the captured image and the user input information, obtaining analysis results and presenting the analysis results.

[0148] In one embodiment, when the processor executes the computer program, it further performs the following steps: in response to a trigger operation of the information acquisition control, acquiring a target image from various captured images; and acquiring analysis results based on the target image and information input by the user.

[0149] In one embodiment, when the processor executes the computer program, it further performs the following steps: in response to a triggering operation of the information acquisition control, determining the captured image corresponding to the triggering of the information acquisition control from among the captured images, and acquiring the target image based on the captured image at the triggering of the information acquisition control.

[0150] In one embodiment, when the processor executes the computer program, it further performs the following steps: displaying the target image on a first interface during the process of collecting information input by the user.

[0151] In one embodiment, the processor, when executing a computer program, further performs the following steps: in response to a trigger operation of the voice acquisition control, acquires target voice input by the user via a microphone.

[0152] In one embodiment, when the processor executes the computer program, it further performs the following steps: in response to a trigger operation of the voice acquisition control, it acquires a candidate signal input by the user through a microphone; it performs voice activity detection on the candidate signal, and if the candidate signal is determined to be a voice signal, it uses the candidate signal as the target voice input by the user.

[0153] In one embodiment, when the processor executes the computer program, it further performs the following steps: performing optical character recognition on the captured image to obtain captured image text; and obtaining analysis results based on the captured image, captured image text, and user input information.

[0154] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps: in response to a trigger operation, displaying a first interface and displaying an image captured by a camera on the first interface; the first interface further includes an information acquisition control; in response to a trigger operation on the information acquisition control, acquiring information input by a user; and based on the captured image and the information input by the user, obtaining analysis results and presenting the analysis results.

[0155] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a trigger operation of the information acquisition control, acquiring a target image from various captured images; and acquiring analysis results based on the target image and information input by the user.

[0156] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a triggering operation of the information acquisition control, determining from the various captured images the captured image corresponding to the triggering of the information acquisition control, and acquiring the target image based on the captured image at the triggering of the information acquisition control.

[0157] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: displaying the target image on the first interface during the process of collecting information input by the user.

[0158] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a trigger operation of the voice acquisition control, acquiring the target voice input by the user via a microphone.

[0159] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a trigger operation of the voice acquisition control, acquiring candidate signals input by the user through a microphone; performing voice activity detection on the candidate signals; and if the candidate signals are determined to be voice signals, using the candidate signals as the target voice input by the user.

[0160] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: performing optical character recognition on the captured image to obtain captured image text; and obtaining analysis results based on the captured image, captured image text, and user input information.

[0161] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps: in response to a trigger operation, displays a first interface and displays a captured image obtained by a camera on the first interface; the first interface further includes an information acquisition control; in response to a trigger operation on the information acquisition control, acquires information input by the user; and based on the captured image and the user input information, obtains analysis results and presents the analysis results.

[0162] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a trigger operation of the information acquisition control, acquiring a target image from various captured images; and acquiring analysis results based on the target image and information input by the user.

[0163] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a triggering operation of the information acquisition control, determining from the various captured images the captured image corresponding to the triggering of the information acquisition control, and acquiring the target image based on the captured image at the triggering of the information acquisition control.

[0164] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: displaying the target image on the first interface during the process of collecting information input by the user.

[0165] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a trigger operation of the voice acquisition control, acquiring the target voice input by the user via a microphone.

[0166] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to a trigger operation of the voice acquisition control, acquiring candidate signals input by the user through a microphone; performing voice activity detection on the candidate signals; and if the candidate signals are determined to be voice signals, using the candidate signals as the target voice input by the user.

[0167] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: performing optical character recognition on the captured image to obtain captured image text; and obtaining analysis results based on the captured image, captured image text, and user input information.

[0168] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0170] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0171] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An interactive processing method, characterized in that, The method includes: In response to a trigger operation, a first interface is displayed, and the captured image obtained through the camera is displayed on the first interface; the first interface also includes an information collection control; In response to a trigger operation on the information collection control, information input by the user is collected; Based on the captured images and the information input by the user, the analysis results are obtained and presented.

2. The method according to claim 1, characterized in that, The method further includes: In response to a trigger operation of the information acquisition control, a target image is acquired from each of the captured images; The step of obtaining analysis results based on the captured image and the information input by the user includes: The analysis results are obtained based on the target image and the information input by the user.

3. The method according to claim 2, characterized in that, The step of acquiring a target image from each of the captured images in response to a trigger operation of the information acquisition control includes: In response to the triggering operation of the information acquisition control, the captured image corresponding to the triggering of the information acquisition control is determined from each of the captured images, and the target image is obtained based on the captured image at the triggering of the information acquisition control.

4. The method according to claim 2, characterized in that, The method further includes: During the process of collecting user input information, the target image is displayed on the first interface.

5. The method according to claim 1, characterized in that, The information acquisition control includes a voice acquisition control; the step of acquiring user input information in response to a trigger operation on the information acquisition control includes: In response to a trigger operation on the voice acquisition control, the target voice input by the user is acquired via the microphone.

6. The method according to claim 5, characterized in that, The step of responding to a trigger operation on the voice acquisition control by acquiring the target voice input by the user through the microphone includes: In response to a trigger operation on the voice acquisition control, candidate signals input by the user are acquired via the microphone; The candidate signal is subjected to speech activity detection. If the candidate signal is determined to be a speech signal, then the candidate signal is used as the target speech input by the user.

7. The method according to any one of claims 1 to 6, characterized in that, After displaying the captured image via camera on the first interface, the method further includes: Optical character recognition is performed on the captured image to obtain the captured image text; The step of obtaining analysis results based on the captured image and the information input by the user includes: The analysis results are obtained based on the captured image, the captured image text, and the information input by the user.

8. An interactive processing device, characterized in that, The device includes: The first interface display module is used to respond to a trigger operation, display a first interface, and display the captured image obtained by the camera on the first interface; the first interface also includes an information collection control; The user input information collection module is used to collect user input information in response to the trigger operation of the information collection control; The analysis result acquisition module is used to acquire analysis results based on the captured image and the information input by the user, and to present the analysis results.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.