Interaction method, apparatus, device, and program medium for article identification

By inputting the target video into the interactive interface and using an intelligent agent for recognition, the inefficiency caused by switching between multiple applications is solved, resulting in a more efficient and better user experience.

CN122132599APending Publication Date: 2026-06-02BEIJING WODONG TIANJUN INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
Filing Date
2026-03-12
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, users need to frequently switch between multiple applications to identify items in a video, resulting in low operational efficiency and a poor user experience.

Method used

By inputting the target video into the interactive interface, the intelligent agent performs recognition and displays a summary of the features of the object to be recognized, avoiding frequent switching of applications.

Benefits of technology

It improves the accuracy and efficiency of object recognition, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132599A_ABST
    Figure CN122132599A_ABST
Patent Text Reader

Abstract

This disclosure provides an interactive method, apparatus, device, and program medium for object recognition, relating to the field of artificial intelligence technology, and more specifically, to the fields of computer vision, multimodal recognition, and intelligent interaction technology. The interactive method for object recognition includes: displaying a target video of an object to be recognized input via an interactive interface; and using an intelligent agent to recognize the target video of the object to be recognized, and displaying a first recognition result on the interactive interface; wherein the first recognition result includes feature summary information of the object to be recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to the fields of computer vision, multimodal recognition, and intelligent interaction technology, specifically to an interactive method, apparatus, device, and program medium for object recognition. Background Technology

[0002] With the development of artificial intelligence technology, it can identify objects based on images or recommend similar items based on the identification results.

[0003] In realizing the present invention, the inventors discovered that the related technology has at least the following problems: when a user needs to identify objects in a video, they need to use other applications to take screenshots of the video frame by frame before they can identify the objects. This interactive method, which involves frequent switching between multiple applications, has low operating efficiency and a poor user experience. Summary of the Invention

[0004] In view of this, the present disclosure provides an interactive method for object recognition, comprising: displaying a target video of an object to be recognized input via an interactive interface; and using an intelligent agent to recognize the target video of the object to be recognized, and displaying a first recognition result on the interactive interface; wherein the first recognition result includes feature summary information of the object to be recognized.

[0005] One aspect of this disclosure provides an interactive device for object recognition, comprising: a display module and a display module.

[0006] The display module is used to display the target video of the object to be identified, input via the interactive interface. The presentation module is used to utilize an intelligent agent to identify the target video of the object to be identified and to display the first identification result on the interactive interface; wherein the first identification result includes feature summary information of the object to be identified.

[0007] Another aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the methods described above.

[0008] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the methods described above.

[0009] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the methods described above.

[0010] According to embodiments of this disclosure, a target video of an object to be identified is input via an interactive interface. An intelligent agent identifies the target video and displays the identification result, representing feature summary information of the object, on the interactive interface. Since video provides richer object information than a single image, the accuracy of the object identification result can be further improved. This at least partially overcomes the technical problem of low operational efficiency caused by frequent switching between multiple applications, thereby achieving the technical effect of improving operational efficiency and user experience. Attached Figure Description

[0011] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0012] Figure 1 The illustration schematically shows an exemplary system architecture to which the interactive methods and apparatus for object recognition disclosed herein can be applied;

[0013] Figure 2 A flowchart illustrating an interactive method for item recognition according to an embodiment of the present disclosure is shown schematically.

[0014] Figure 3A The diagram illustrates an interactive interface of an interactive method for item recognition according to an embodiment of the present disclosure.

[0015] Figure 3B The diagram illustrates an interactive interface of an interactive method for item recognition according to another embodiment of the present disclosure.

[0016] Figure 4 The illustration shows an interactive interface diagram in a scene with multiple items to be identified in a target video according to an embodiment of the present disclosure;

[0017] Figure 5A The illustration shows a schematic diagram of an interface for interacting with a playback progress control according to an embodiment of the present disclosure;

[0018] Figure 5B The illustration shows a schematic diagram of an interface for interacting with a playback progress control according to another embodiment of the present disclosure;

[0019] Figure 6 The illustration shows a schematic diagram of an interface for interacting with an item location marker control according to an embodiment of the present disclosure;

[0020] Figure 7 A flowchart illustrating a process for processing a target video to obtain a recognition result, according to an embodiment of the present disclosure, is shown schematically.

[0021] Figure 8A The illustration shows a schematic diagram of an interactive interface for object recognition combining video and voice according to an embodiment of the present disclosure;

[0022] Figure 8B The illustration shows a schematic diagram of an interactive interface for object recognition combining video and voice according to another embodiment of the present disclosure;

[0023] Figure 8C The illustration shows a schematic diagram of an interactive interface for object recognition combining video and voice according to yet another embodiment of the present disclosure;

[0024] Figure 9 A schematic flowchart illustrating a process for processing target video and audio to obtain recognition results according to an embodiment of the present disclosure is shown.

[0025] Figure 10 A block diagram of an interactive device for item recognition according to embodiments of the present disclosure is schematically shown; and

[0026] Figure 11 A block diagram of an electronic device 1100 suitable for implementing a robot according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0027] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0028] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0030] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0031] Currently, users can identify interesting objects in their daily lives using mobile devices. However, when users are in a confined space and take pictures of large objects, they can only capture a part of the object. Therefore, it is impossible to extract all the object features based on a single image, resulting in low accuracy of the recognition results.

[0032] When users want to identify objects in videos they have previously filmed or downloaded, they need to take screenshots frame by frame of the video and then upload the screenshots, including the objects, to an application with image recognition capabilities for identification. This often requires frequent switching between multiple applications, resulting in low efficiency and a poor user experience.

[0033] In view of this, the present disclosure provides an interactive method for object recognition, comprising: displaying a target video of an object to be recognized input via an interactive interface; and using an intelligent agent to recognize the target video of the object to be recognized, and displaying a first recognition result on the interactive interface; wherein the first recognition result includes feature summary information of the object to be recognized.

[0034] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0035] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.

[0036] Figure 1 An exemplary system architecture 100 is schematically illustrated, to which the interactive methods and apparatus for item recognition disclosed herein can be applied. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0037] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0038] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, intelligent recognition applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).

[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, smartwatches, etc.

[0040] Server 105 can be a server providing various services, such as a backend management server supporting websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. For example, server 105 can process and analyze captured target videos to obtain object recognition results.

[0041] It should be noted that the interactive method for item recognition provided in this embodiment can be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103. Accordingly, the interactive device for item recognition provided in this embodiment can also be disposed in the first terminal device 101, the second terminal device 102, and the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0042] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0043] Figure 2 A flowchart illustrating an interactive method for item recognition according to an embodiment of the present disclosure is shown schematically.

[0044] like Figure 2 As shown, the method includes operations S210~S220.

[0045] When operating S210, the target video of the object to be identified is displayed, which is input via the interactive interface.

[0046] When operating the S220, the intelligent agent is used to identify the target video of the object to be identified, and the first identification result is displayed on the interactive interface.

[0047] In some embodiments, users can upload target videos via a network to complete the input operation. In some embodiments, when a user is taking a picture of an object to be identified, the intelligent agent can be automatically triggered to perform a recognition operation on the target video when the filming operation is terminated.

[0048] The action to stop recording an object can refer to the operation of a recording control on the terminal device used to record the object. This recording control can be a physical control configured on the terminal device or a virtual switch control configured on the terminal device's display screen.

[0049] For example, clicking the virtual switch control during shooting can stop the shooting operation. Another example is when a user presses and holds the virtual switch control while continuously shooting; when the user removes their hand from the virtual switch control, shooting stops. Therefore, stopping shooting can be achieved by changing the user's hand from a contact state to a non-release state with the virtual switch control.

[0050] In some embodiments, the first identification result may include feature summary information of the item to be identified, such as the selling points information of the item to be identified.

[0051] In addition to directly displaying the first recognition result on the interactive interface, in some embodiments, since the time required for the intelligent agent to recognize the target video varies, the first operation status information can be displayed on the display interface of the target video captured for the object to be recognized using a first display control. The first operation status information indicates the progress of the first processing operation required to process the target video to obtain the recognition result. This allows the user to understand the video processing progress in real time, enriches the interface content during video processing, and provides the user with multiple visual interface options, further enhancing the user experience.

[0052] The first processing operation may include extracting video frames, recognizing video frames, retrieving item information, analyzing search results, and so on.

[0053] The status information of the first operation can be dynamically updated whenever the progress of any operation changes during the first processing operation. For example, at time t, the status information indicates that video frames are being extracted; at time t+1, the status information could indicate that video frames have been extracted, or that video frame recognition is in progress. Users can monitor the video processing progress in real time, enriching the interface content during video processing.

[0054] During the processing of the target video, the target video display interface remains visible to the user. Since video processing and object recognition take time, displaying the initial operation status information in the form of a control provides users with multiple visual interface options, further enriching the user experience. For example, when a user does not wish to view the processing progress, they can move the display control to the bottom of the screen to view the captured target video.

[0055] Once the first processing operation has been completed, the first operation status information is switched to the recognition result. This recognition result may include not only the identified item attribute information, but also information designed to increase the target audience's interest in the identified item.

[0056] For example, the item to be identified could be a refrigerator. In related examples, the identification results are relatively simple, only including item attribute information, such as: brand A refrigerator, refrigerator model, etc. Alternatively, it could include recommendation information for similar items retrieved based on the identification results, such as: brand B refrigerator, etc.

[0057] In this embodiment, item description information from an external knowledge base can be retrieved based on item attribute information. This external knowledge base can be constructed from publicly available information describing various item characteristics, such as information from search platforms or item trading platforms. Then, utilizing the natural language understanding capabilities of a large language model, information is generated based on the item description information to increase the target audience's interest in the identified item. Increasing the target audience's interest can be achieved by focusing on the item's inherent characteristics, such as the advantages of Brand A refrigerators compared to similar items from other brands. Alternatively, it can be based on user transaction preferences, such as the ability of Brand A refrigerators to support whole-house smart home control.

[0058] This embodiment of the disclosure utilizes an intelligent agent to identify a target video of an object to be identified, input via an interactive interface, and displays the identification result, which is a feature summary information representing the object, on the interactive interface. Since video provides richer object information than a single image, it not only further improves the accuracy of object identification results but also avoids frequent switching between multiple applications, thus improving operational efficiency.

[0059] Figure 3A The diagram illustrates an interactive interface of an interactive method for item recognition according to an embodiment of the present disclosure.

[0060] like Figure 3A As shown, the user is recording a video of a kitchen, which may contain items such as a refrigerator, cabinets, and sink. Recording stops when the user's hand leaves the virtual switch control. The captured target video is then input into the interactive interface. The intelligent agent recognizes the target video and displays the first recognition result on the target video display interface 310 using the first display control 320.

[0061] Figure 3B The diagram illustrates an interactive interface of an interactive method for item recognition according to another embodiment of the present disclosure.

[0062] like Figure 3B As shown, the user is recording a video of a kitchen, which may contain items such as a refrigerator, cabinets, and sink. Recording stops when the user's hand leaves the virtual switch control. The captured target video is then input into the interactive interface. The intelligent agent recognizes the target video and displays the first operation status information on the target video display interface 310 using the first display control 320.

[0063] like Figure 3B As shown, the first operation status information indicates that the operations of extracting key images from video frames and identifying the main objects in video frames have been completed, and the object analysis operation is currently being performed.

[0064] When the first processing operation has been completed, the information in the first display control 320 changes from the first operation status information to the first recognition result. For example... Figure 3B As shown, the identification result may include the item attribute "refrigerator" and the advantages of the item to be identified, such as "whole-house smart control" and "intelligent preservation".

[0065] In real-world scenarios, users may want to identify multiple items simultaneously; therefore, the items in different frames of the captured target video are not exactly the same.

[0066] When there are multiple items to be identified, the intelligent agent is used to identify the target video of the items to be identified and the first identification result is displayed on the interactive interface. This may include the following operations: using the intelligent agent to identify the target video of the items to be identified and displaying the item image controls of each of the multiple items to be identified; and in response to a trigger operation on any item image control among the multiple item image controls, determining the first identification result of the first item corresponding to any item image control from the first identification result and displaying the identification result of the first item.

[0067] Figure 4 The illustration shows an interactive interface diagram of a target video with multiple objects to be identified according to an embodiment of the present disclosure.

[0068] like Figure 4 As shown, the items to be identified may include a refrigerator and a cabinet. An item image control for a "refrigerator" and an item image control for a "cabinet" are displayed using a first display control 320.

[0069] When there are multiple items to be identified, embodiments of this disclosure can determine the display positions of multiple item image controls based on the sorting results of the display ratios of the multiple items to be identified in the target video.

[0070] For example, in the target video, the display ratio of the "refrigerator" in each video frame is greater than that of the "cabinet". This display ratio can be the ratio of the display area of ​​the "refrigerator" and the "cabinet" in the same frame, or it can be the ratio of the number of video frames containing the "refrigerator" to the number of video frames containing the "cabinet" in the target video. This embodiment of the disclosure does not specifically limit this. Therefore, the display position of the item image control of the "refrigerator" is before the display position of the item image control of the "cabinet".

[0071] like Figure 4 As shown, the item image control for "refrigerator" is selected by a box, and the first display control 320 displays the recognition result of the refrigerator.

[0072] When the target subject takes a picture of the item to be identified, the item of interest is usually photographed for a longer time or at close range. Therefore, the identification results of the item with the largest display ratio are displayed first, which is more in line with the user's identification intention.

[0073] In this embodiment of the disclosure, in response to a selection operation of a target item image control among a plurality of item image controls, the recognition result of the first item is switched to the recognition result of the second item corresponding to the first target item image control.

[0074] like Figure 4 As shown, when a user wants to view the recognition result of the "cabinet", they can perform a selection operation on the "cabinet" item image control to switch the information in the first display control 320 from the refrigerator recognition result to the cabinet recognition result.

[0075] When the target video contains multiple items to be identified, the identification results of the multiple items are displayed and switched by switching the corresponding item image control. This allows users to repeatedly view the identification results of different items without having to perform secondary identification, saving the waiting time for secondary identification and further improving the user experience.

[0076] The operation status and recognition results are displayed in the form of display controls, which facilitates users to switch between the display interface and display controls according to actual needs. The above method may also include the following operations: in response to a trigger operation of the playback progress control on the interactive interface for the target playback time of the target video, displaying the target image in the target video corresponding to the target playback time; and determining a second recognition result from the first recognition result corresponding to the second item in the target image, and displaying the second recognition result.

[0077] Figure 5A The illustration shows a schematic diagram of an interface for interacting with a playback progress control according to an embodiment of the present disclosure.

[0078] like Figure 5A As shown, when the user moves the first display control 320 along the direction of the arrow, the first display control 320 is moved to the bottom of the display interface. A preview image of the target video is displayed on the display interface 310. This preview image can be the first frame of the target video, or it can be the first keyframe among multiple keyframes extracted from the target video.

[0079] The display interface is also equipped with a playback progress control 520 for controlling the playback of the target video, so that users can play the target video by operating the playback progress control 520.

[0080] Users can also select a target image from the target video by operating the playback progress control 520. For example, when a user drags the virtual touch control on the playback progress control 520 to the target position, it indicates that the target playback time corresponding to the target position has been selected. The target image corresponding to the target playback time is then displayed at that target position using the second display control 510.

[0081] like Figure 5A As shown, the target image can be an image with a range hood as the main object. At this time, the information in the first display control 320 can be switched from the recognition result of "refrigerator" to the recognition result of "range hood".

[0082] By manipulating the playback progress control, users can select the image corresponding to a specific moment for targeted recognition of the target object, further satisfying users' personalized needs for object recognition and improving the user experience.

[0083] During the processing of the target video, it's possible to miss identifying items that are displayed in a small proportion of the video. Examples include switches above the countertop or kitchen paper towels in a corner.

[0084] When the user determines that the target image includes the missed item by operating the progress control, the recognition operation for the target image can be restarted.

[0085] According to embodiments of this disclosure, the above method may further include the following operations: in response to determining that the first recognition result does not include the recognition result of the second item, using an intelligent agent to recognize the target image and displaying the second recognition result of the second item.

[0086] Whether an item in the target image is in the recognition result obtained based on the first processing operation can be determined based on the correlation between the recognition result and the item in the target image.

[0087] For example, a predetermined correlation threshold can be pre-configured. When the correlation between the recognition result and the items in the target image is greater than the predetermined correlation threshold, it means that the recognition result includes the recognition result of the items in the target image, and no recognition operation is required for the target image. When the correlation between the recognition result and the items in the target image is less than or equal to the predetermined correlation threshold, it means that the recognition result does not include the recognition result of the items in the target image, and recognition operation for the target image needs to be performed.

[0088] Figure 5B The diagram illustrates an interface for interacting with a playback progress control according to another embodiment of the present disclosure.

[0089] like Figure 5B As shown, since the recognition result of the main item "kitchen paper towel" in the target image is not in the recognition result obtained based on the first processing operation, the recognition result is switched to the second operation status information. At this time, the second operation status information is displayed in the first display control 320.

[0090] The second operation status information indicates the progress of the second processing operation required to process the target image to obtain the recognition result of the third item in the target image.

[0091] like Figure 5B As shown, the second processing operation may include identifying the main object in the image and analyzing the object. The current operation progress indicates that the operation of identifying the main object in the image has been completed, and the analysis operation for the main object is being performed.

[0092] Once the second processing operation has been completed, the information in the first display control 320 changes from the second operation status information to the recognition result of "kitchen paper towels". This recognition result also includes information to increase the target object's interest in "kitchen paper towels".

[0093] In some embodiments, when an item in a specific image selected by a user needs to be re-identified, an intelligent agent can be used to identify the item in the specific image and display the identification result on the interactive interface to meet the user's identification needs for a specific item.

[0094] In some embodiments, when an item in a specific image selected by the user needs to be re-identified, the information displayed in the control can be switched to information representing the progress of the re-identification operation. This avoids the problem of the control still displaying the initial recognition result during re-identification, which could mislead the user into believing that a recognition error has occurred, thus degrading the user experience. Furthermore, after the re-identification operation is completed, the recognition result of the item in the target image is promptly switched to further meet the user's need for recognition efficiency.

[0095] For target videos taken by users, items that are displayed in a proportion greater than a predetermined display proportion threshold in each video frame are usually prioritized for identification. Therefore, there may be cases where items are missed.

[0096] In view of this, the interactive method provided in the embodiments of this disclosure supports users to move the item location marker control so as to select any item for identification.

[0097] According to embodiments of this disclosure, in response to a movement operation of a display control on an interactive interface for displaying a first recognition result, a preview image of the target video is displayed.

[0098] By moving the display controls, a preview image of the target video is displayed, allowing users to specify the items they want to identify in the preview image, further meeting their personalized needs.

[0099] In some embodiments, the above method may further include the following operations: in response to a movement operation of moving from the current position to the target position for the item location identifier control in the interactive interface, identifying the item displayed at the target position in the preview image as a third item; and using an agent to identify the preview image and displaying the third identification result of the third item on the interactive interface.

[0100] Figure 6 The diagram illustrates an interface for interacting with an item location marker control according to an embodiment of the present disclosure.

[0101] like Figure 6 As shown in the preview image, the initial item selected by the item location indicator control 610 is a "refrigerator". The user can move the item location indicator control 610, for example, from the location of the "refrigerator" to the location of the "faucet", and scale the size of the item location indicator control 610 to match the display area of ​​the "faucet". In some embodiments, after moving the item location indicator control 610, the size of the item location indicator control 610 can also adaptively adjust according to the display area of ​​the item at the new location.

[0102] For items at new locations that require re-identification, the third operation status information can be displayed in the first display control 320. Alternatively, the initial identification result in the first display control 320 can be switched to the third operation status information. Or, the third identification result can be displayed directly without showing the third operation status information.

[0103] The third operation status information indicates the progress of the third processing operation required to process the preview image in order to identify the fourth item. For example... Figure 6 As shown, the third processing operation may include the operation of recognizing the main object in the image and the operation of analyzing the main object.

[0104] It should be noted that the main object in the image at this time is the "faucet" selected by the moved object position marker control 610.

[0105] Therefore, when the third processing operation is completed, the information in the first display control 320 is switched to the recognition result of "faucet".

[0106] When a user reselects the identified item using the item location marker control, the information displayed within the control switches to information indicating the progress of the re-identification operation. This avoids the problem of the control still showing the initial identification result during re-identification, which could mislead users into believing there was an error and degrade the user experience. Furthermore, after the re-identification operation is complete, the identification result of the item in the target image is promptly switched, further satisfying the user's demand for recognition efficiency.

[0107] The following is combined Figure 7 The processing procedure for the target video is explained in detail.

[0108] According to embodiments of this disclosure, using an intelligent agent to identify a target video of an object to be identified may include the following operations: performing object detection on multiple target frames in the target video to obtain object detection results for each of the multiple target frames; performing correlation aggregation on the object detection results for each of the multiple target frames to obtain an aggregation result; wherein the aggregation result includes an image of the object to be identified extracted from the multiple target frames based on the object detection results; and inputting the object description information associated with the aggregation result and the image of the object to be identified into a multimodal large model to output a first identification result.

[0109] Figure 7 A schematic flowchart illustrating a process for processing a target video to obtain a recognition result according to an embodiment of the present disclosure is shown.

[0110] like Figure 7 As shown, firstly, N target frames can be extracted from the target video 701, where N is an integer greater than 1. The target frames can be key frames in the target video 701 that differ significantly from adjacent frames.

[0111] Then, an object detection model can be used to perform object detection on each target frame, resulting in an item detection result 703 for each frame. This item detection result can include the item's position and attributes in each target frame.

[0112] Next, the object detection results of N target frames can be correlated and aggregated. This can be understood as identifying objects with a correlation or similarity greater than a predetermined threshold as the same object, thus obtaining the aggregation result 704. When the target video includes multiple objects to be identified, the aggregation result 704 gathers the images of the objects to be identified extracted from multiple target frames based on the object detection results, i.e., the object slice 706.

[0113] Then, based on the image of the item to be identified, the item description information 705 of the item to be identified can be retrieved from the knowledge base.

[0114] Finally, the retrieved item description information 705 and item image 706 are input into the multimodal large model to generate the recognition result 707.

[0115] For the image recognition operation in this embodiment of the disclosure, only the frame extraction operation of the video is removed. The remaining operations are similar to the video recognition operation and will not be described in detail here.

[0116] According to embodiments of this disclosure, by performing target detection on multiple target frames in a target video separately and then aggregating the detection results, the problem of inaccurate recognition results caused by the lack of information provided by a single image when relying on images that only capture local features of the object to be identified is at least solved, thus further improving the accuracy of object recognition. Simultaneously, by utilizing a multimodal large model based on semantic understanding of image and object description information, recognition results are generated to enhance the target object's interest in the object, further improving the user experience compared to methods in related examples that can only identify object attribute information.

[0117] The examples provided in the relevant samples contained only basic item attribute information or recommendations for similar items, which is insufficient to meet users' personalized needs.

[0118] Based on object recognition using target video, this embodiment of the present disclosure can also incorporate voice commands input by the user during video recording based on their own needs, so that the final recognition result can meet the personalized needs of different users.

[0119] Therefore, the method of this disclosure embodiment may further include the following operations: during the process of the target object performing a shooting operation on the object to be identified, displaying guidance information related to the object to be identified on the interactive interface; wherein, the guidance information is used to guide the target object to input expected voice information during the shooting operation; using an intelligent agent to identify the target video and the actual audio received during the shooting, and displaying a fourth identification result on the interactive interface.

[0120] According to embodiments of this disclosure, guidance information is used to guide the target object to input expected voice information during the shooting operation. This expected voice information can be obtained in real-time by matching attribute information of the object to be identified obtained from the target video. For example, multiple expected voice information related to different objects can be pre-configured, and then the corresponding expected voice information can be dynamically matched based on the identified attribute information. In some embodiments, it can also be dynamically matched based on user historical preferences, or the expected voice information can be dynamically generated based on the attribute information of the object to be identified combined with user historical preferences.

[0121] For example, if the item to be identified is a "sweater", the expected voice message could be "You can try saying what season this sweater is suitable for?", or if the user's historical preferences are related to clothing matching, the dynamically matched expected voice message could be "You can try saying how to match this sweater?"

[0122] The purpose of this guidance information is to guide users in inputting their own needs so that the multimodal large model can output recognition results that match those needs. However, it does not restrict users from inputting voice information that is the same as or related to the guidance information. Regardless of whether the actual voice information input by the user is related to the expected voice information, the multimodal large model will ultimately output recognition results based on the actual voice information.

[0123] When the user inputs voice, the fourth operation status information displayed by the first display control indicates the progress of the fourth processing operation required to process the target video and the actual audio received during shooting to obtain the recognition result.

[0124] Figure 8A The illustration shows a schematic diagram of an interactive interface for object recognition combining video and voice according to an embodiment of the present disclosure.

[0125] like Figure 8AAs shown, when a user is taking a picture of a cashmere coat, the shooting interface 810 displays the guidance message "Try saying: How to match this coat to look good" 801, in addition to the captured image. This guidance message can also serve as a prompt to the user that the current application supports voice input, so that the user can input voice commands according to their actual needs during the shooting process.

[0126] At this point, the actual audio input by the user could be "What pants would look good with this coat?" 802. It can be seen that the semantics of the user's actual audio input are similar to the guiding information, and specifically indicate that the type of item to be paired is "pants".

[0127] Next, when the user stops shooting, the fourth operation status information is displayed in the first display control 320.

[0128] like Figure 8A As shown, the fourth processing operation may include extracting key video frames, reading audio content, identifying main objects in the video, and analyzing objects and questions. The "question" refers to the question mentioned in the actual audio content input by the user.

[0129] Once the fourth processing operation is completed, the fourth operation status information in the first display control 320 switches to the recognition result. This recognition result includes information from the user's actual voice response, such as "This cashmere garment is recommended to be paired with...". The matching suggestions in the recognition result are all related to the specific item type "trousers" specified in the actual voice information. Simultaneously, similar matching item recommendations can also be displayed, such as information on other trousers of the same type as "XX trousers" or "YY trousers".

[0130] Figure 8B The diagram illustrates an interactive interface for object recognition combining video and voice according to another embodiment of the present disclosure.

[0131] like Figure 8B As shown, the shooting scene in this embodiment is similar to... Figure 8A The shooting scene and guidance information shown are the same, but the actual voice information entered by the user is "How is the weather today?", which means that the actual voice information entered by the user has a low correlation with the item to be identified, "cashmere coat".

[0132] Therefore, after the fourth processing operation is completed, the final recognition result displayed in the first display control is the recognition result of the item to be recognized, and it is not closely related to the actual voice information. This recognition result is the same as the processing result based solely on video, and may include attribute information of the item to be recognized, such as "cashmere coat," and may also include information to increase the target audience's interest in the item to be recognized, such as "soft and skin-friendly," "lightweight and warm," etc.

[0133] Similarly, you can also display recommendations for similar items, except that the recommended items will be "cashmere coats".

[0134] Figure 8C The diagram illustrates an interactive interface for object recognition combining video and voice according to yet another embodiment of the present disclosure.

[0135] like Figure 8C As shown, the shooting scene in this embodiment is similar to... Figure 8A , Figure 8B The shooting scene and guidance information shown are the same, but the actual voice information entered by the user is "What season is this coat suitable for?", which shows that the actual voice information entered by the user is highly relevant to the item to be identified, "cashmere coat".

[0136] Therefore, after the fourth processing operation has been completed, the final recognition result displayed in the first display control is the information used to answer "What season is this coat suitable for?": "This cashmere coat is suitable for wearing in the early spring season when the temperature is 0~5℃".

[0137] At this point, the recommended item type in the similar item recommendations can be "cashmere coat".

[0138] Building upon the video, guidance information has been added to encourage users to input their needs via voice, thereby enabling the output of recognition results that match the user's personalized requirements and further enhancing the user's interactive experience.

[0139] In real-world applications, the audio input by users during the recording process may contain redundant content or have low relevance to the object to be identified. If this content is input into a multimodal model, it may mislead the model into outputting results that deviate from the user's actual intent.

[0140] Therefore, using an intelligent agent to identify the target video and the actual audio received during filming can include the following operations: performing speech recognition on the actual audio to obtain actual speech information; and using a multimodal large model to extract the actual speech information to obtain a speech summary; wherein the speech summary indicates the core semantics of the actual speech information.

[0141] The following is combined Figure 9 The processing flow for obtaining recognition results from target video and audio is described in detail.

[0142] Figure 9 A schematic flowchart illustrating a process for processing target video and audio to obtain recognition results according to an embodiment of the present disclosure is shown.

[0143] like Figure 9As shown, firstly, N target frames can be extracted from the target video 701, where N is an integer greater than 1. The target frames can be key frames in the target video 701 that differ significantly from adjacent frames.

[0144] Then, an object detection model can be used to perform object detection on each target frame, resulting in an item detection result 703 for each frame. This item detection result can include the item's position and attributes in each target frame.

[0145] The identification of target video and actual audio received during recording using an intelligent agent may further include the following operations: performing target detection on multiple target frames in the target video to obtain the item detection results for each of the multiple target frames; performing correlation aggregation on the item detection results for each of the multiple target frames to obtain an aggregation result; wherein, the aggregation result includes images of the items to be identified extracted from the multiple target frames based on the item detection results; in response to determining that the correlation between the speech summary and the item to be identified is greater than a predetermined correlation threshold, performing intent recognition on the speech summary using a multimodal large model to obtain the questioning intent of the target object regarding the item to be identified; and inputting item description information related to the questioning intent, item description information associated with the aggregation result, and the image of the item to be identified into the multimodal large model, and outputting the recognition result; the recognition result also includes item information related to the questioning intent.

[0146] For example, the correlation between the item detection results of N target frames can be aggregated. This can be understood as identifying items whose correlation or similarity is greater than a predetermined threshold as the same item, thus obtaining the aggregation result 704. When the target video includes multiple items to be identified, the aggregation result 704 gathers the images of the items to be identified extracted from multiple target frames based on the item detection results, i.e., the item slice 706.

[0147] Then, based on the image of the item to be identified, the item description information 705 of the item to be identified can be retrieved from the knowledge base.

[0148] Simultaneously, speech recognition is performed on the actual audio 910 to obtain speech information 920. Then, core semantic extraction is performed on the speech information 920 to obtain speech summary 930. Intent recognition is then performed on the speech summary 930 to obtain the questioning intent 940.

[0149] Finally, the retrieved item description information 705, item image 706, and question intent 940 are input into the multimodal large model to generate recognition result 940.

[0150] The multimodal large model determines the recognition result by judging the relevance between the speech summary and the item to be recognized. When the relevance between the speech summary and the item to be recognized is greater than a predetermined relevance threshold, the output recognition result includes information related to the question's intent (such as...). Figure 8A and 8C As shown). When the relevance between the speech summary and the item to be identified is less than or equal to a predetermined relevance threshold, the output recognition result 950 includes the feature summary information of the item to be identified (such as...). Figure 8B (As shown).

[0151] In some embodiments, speech summarization indicates the core semantics of the actual speech information. By extracting the core speech of the actual speech information, the interference of redundant information in the actual speech information on the multimodal large model analysis process can be reduced, the amount of input data of the multimodal large model can be reduced, and the interaction efficiency can be further improved.

[0152] Figure 10 A block diagram of an interactive device for item recognition according to an embodiment of the present disclosure is shown schematically.

[0153] like Figure 10 As shown, the interactive device 1000 for object recognition may include a display module 1010 and a display module 1020.

[0154] Display module 1010 is used to display the target video of the object to be identified, input via the interactive interface.

[0155] The display module 1020 is used to identify a target video of an object to be identified using an intelligent agent, and to display a first identification result on an interactive interface; wherein the first identification result includes feature summary information of the object to be identified.

[0156] According to embodiments of this disclosure, the items to be identified include multiple items; the display module may include a first display submodule and a second display submodule.

[0157] The first display submodule is used to use an intelligent agent to identify the target video of the object to be identified and to display the object image controls of multiple objects to be identified.

[0158] The second display submodule is used to respond to a trigger operation for any of the multiple item image controls, determine the first recognition result of the first item corresponding to any item image control from the first recognition result, and display the recognition result of the first item.

[0159] According to embodiments of this disclosure, the display module may further include: a third display submodule and a fourth display submodule.

[0160] The third display submodule is used to respond to the trigger operation of the playback progress control on the interactive interface for the target playback time of the target video, and to display the target image in the target video corresponding to the target playback time.

[0161] The fourth display submodule is used to determine the second recognition result corresponding to the second item in the target image from the first recognition result, and to display the second recognition result.

[0162] According to embodiments of this disclosure, the interactive device 1000 for item recognition may further include: a first recognition module, configured to, in response to determining that the first recognition result does not include the recognition result of the second item, use an intelligent agent to recognize the target image and display the second recognition result of the second item.

[0163] According to embodiments of this disclosure, the interactive device 1000 for object recognition may further include: a preview image display module, configured to display a preview image of a target video in response to a movement operation of a display control on an interactive interface for displaying a first recognition result.

[0164] According to embodiments of this disclosure, the interactive device 1000 for item recognition may further include: a determination module and a second recognition module.

[0165] The determination module is used to determine the item displayed at the target location in the preview image as the third item in response to a movement operation of the item location marker control in the interactive interface from the current location to the target location.

[0166] The second recognition module is used to use an intelligent agent to recognize the preview image and display the third recognition result of the third item on the interactive interface.

[0167] According to embodiments of this disclosure, the identification module may include a target detection module, an aggregation module, and a multimodal identification module.

[0168] The target detection module is used to perform target detection on multiple target frames in the target video and obtain the item detection results for each target frame.

[0169] The aggregation module is used to perform correlation aggregation on the item detection results of multiple target frames to obtain an aggregation result; wherein, the aggregation result includes images of the items to be identified extracted from multiple target frames based on the item detection results.

[0170] The multimodal recognition module is used to input the item description information associated with the aggregation result and the image of the item to be recognized into the multimodal large model, and output the first recognition result.

[0171] According to embodiments of this disclosure, the interactive device 1000 for item recognition may further include: a guidance display module and a third recognition module.

[0172] The guidance display module is used to display guidance information related to the object to be identified on the interactive interface during the process of the target object performing a shooting operation on the object to be identified; the guidance information is used to guide the target object to input the expected voice information during the shooting operation.

[0173] The third recognition module is used to identify the target video and the actual audio received during shooting using an intelligent agent, and displays the fourth recognition result on the interactive interface.

[0174] According to embodiments of this disclosure, the third recognition module may include a speech recognition submodule and a multimodal recognition submodule.

[0175] The speech recognition submodule is used to perform speech recognition on actual audio to obtain actual speech information.

[0176] The multimodal recognition submodule is used to extract speech information from actual speech information using a large multimodal model to obtain speech summaries; the speech summaries indicate the core semantics of the actual speech information.

[0177] According to embodiments of this disclosure, the third identification module may further include: a target detection submodule, an aggregation submodule, an intent identification submodule, and a multimodal identification submodule.

[0178] The object detection submodule is used to perform object detection on multiple target frames in the target video and obtain the object detection results for each of the multiple target frames.

[0179] The aggregation submodule is used to perform correlation aggregation on the item detection results of multiple target frames to obtain the aggregation result; wherein, the aggregation result includes the image of the item to be identified extracted from multiple target frames based on the item detection results.

[0180] The intent recognition submodule is used to perform intent recognition on the speech summary in response to determining that the relevance between the speech summary and the item to be identified is greater than a predetermined association threshold, thereby obtaining the target object's questioning intent regarding the item to be identified.

[0181] A multimodal recognition submodule is used to input item description information related to the questioning intent, item description information associated with the aggregation result, and an image of the item to be recognized into a multimodal large model, and output a recognition result; the recognition result also includes item information related to the questioning intent. Any one or more of the modules, submodules, units, and subunits according to embodiments of this disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of this disclosure can be decomposed into multiple modules for implementation. Any one or more of the modules, submodules, units, and subunits according to embodiments of this disclosure can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, and firmware, or in any suitable combination of any one or more of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to the embodiments of this disclosure may be at least partially implemented as computer program modules, which can perform corresponding functions when the computer program modules are run.

[0182] For example, any plurality of the first display module 1010 and the first switching module 1020 can be combined into one module / unit / subunit, or any one of the modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the first display module 1010 and the first switching module 1020 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any one of the three implementation methods or a suitable combination of any of them. Alternatively, at least one of the first display module 1010 and the first switching module 1020 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0183] It should be noted that the interactive device part in the embodiments of this disclosure is related to the interactive method part in the embodiments of this disclosure. The specific description of the interactive device part is referred to the interactive method, and will not be repeated here.

[0184] Figure 11 A block diagram of an electronic device suitable for implementing the methods described above, according to embodiments of the present disclosure, is illustrated schematically. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0185] like Figure 11 As shown, an electronic device 1100 according to an embodiment of the present disclosure includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage portion 1108 into a random access memory (RAM) 1103. The processor 1101 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1101 may also include onboard memory for caching purposes. The processor 1101 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0186] RAM 1103 stores various programs and data required for the operation of electronic device 1100. Processor 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Processor 1101 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 1102 and / or RAM 1103. It should be noted that programs may also be stored in one or more memories other than ROM 1102 and RAM 1103. Processor 1101 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.

[0187] According to embodiments of this disclosure, the electronic device 1100 may further include an input / output (I / O) interface 1105, which is also connected to a bus 1104. The system 1100 may also include one or more of the following components connected to the input / output (I / O) interface 1105: an input section 1106 including a keyboard, mouse, etc.; an output section 1107 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1110 as needed so that computer programs read from it can be installed into the storage section 1108 as needed.

[0188] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1109, and / or installed from removable medium 1111. When the computer program is executed by processor 1101, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0189] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0190] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0191] For example, according to embodiments of this disclosure, a computer-readable storage medium may include one or more memories other than the ROM 1102 and / or RAM 1103 described above and / or ROM 1102 and RAM 1103.

[0192] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.

[0193] When the computer program is executed by the processor 1101, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0194] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1109, and / or installed from the removable medium 1111. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0195] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0196] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0197] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. An interactive method for object recognition, comprising: Display the target video of the object to be identified, input via an interactive interface; as well as The intelligent agent identifies the target video of the object to be identified and displays the first identification result on the interactive interface; wherein, the first identification result includes feature summary information of the object to be identified.

2. The method according to claim 1, wherein, The items to be identified include multiple items; the step of using an intelligent agent to identify target videos of the items to be identified and displaying the first identification result on the interactive interface includes: The system utilizes an intelligent agent to identify target videos of objects to be identified, and displays individual object image controls for multiple objects to be identified; and In response to a trigger operation for any of the plurality of item image controls, a first recognition result for a first item corresponding to any of the item image controls is determined from the first recognition result, and the recognition result for the first item is displayed.

3. The method according to claim 1 or 2, further comprising: In response to a trigger operation on the playback progress control on the interactive interface for the target playback time of the target video, the target image in the target video corresponding to the target playback time is displayed; as well as A second recognition result is determined from the first recognition result and corresponds to the second item in the target image, and the second recognition result is displayed.

4. The method according to claim 3, wherein, The method further includes: In response to determining that the first recognition result does not include the recognition result of the second item, the agent performs recognition on the target image and displays the second recognition result of the second item.

5. The method according to claim 1 or 2, further comprising: In response to a movement operation of a display control on the interactive interface used to display the first recognition result, a preview image of the target video is displayed.

6. The method according to claim 5, further comprising: In response to a movement operation that moves the item location marker control in the interactive interface from its current location to a target location, the item displayed in the preview image at the target location is identified as a third item. as well as The intelligent agent is used to identify the preview image, and the third identification result of the third item is displayed on the interactive interface.

7. The method according to claim 1 or 2, wherein, The process of using an intelligent agent to identify target videos of objects to be identified includes: Target detection is performed on multiple target frames in the target video to obtain the item detection results for each target frame. The object detection results of the multiple target frames are correlated and aggregated to obtain an aggregated result; wherein, the aggregated result includes images of the objects to be identified extracted from the multiple target frames based on the object detection results; and The item description information associated with the aggregation result and the image of the item to be identified are input into the multimodal large model, and the first identification result is output.

8. The method according to claim 1 or 2, further comprising: During the process of the target object taking a picture of the object to be identified, guidance information related to the object to be identified is displayed on the interactive interface; wherein, the guidance information is used to guide the target object to input expected voice information during the process of taking the picture; and The intelligent agent is used to identify the target video and the actual audio received during the shooting, and the fourth identification result is displayed on the interactive interface.

9. The method according to claim 8, wherein the step of using the intelligent agent to identify the target video and the actual audio received during shooting comprises: The actual audio is subjected to speech recognition to obtain actual speech information; as well as The actual speech information is extracted using a multimodal large model to obtain a speech summary; wherein the speech summary indicates the core semantics of the actual speech information.

10. The method according to claim 9, wherein the step of using the intelligent agent to identify the target video and the actual audio received during shooting further comprises: Target detection is performed on multiple target frames in the target video to obtain the item detection results for each target frame. The correlation of the item detection results of each of the multiple target frames is aggregated to obtain an aggregation result; wherein, the aggregation result includes images of the items to be identified extracted from the multiple target frames based on the item detection results; In response to determining that the relevance between the speech summary and the item to be identified is greater than a predetermined association threshold, the multimodal large model is used to perform intent recognition on the speech summary to obtain the target object's questioning intent regarding the item to be identified; and The multimodal large model is input with the item description information related to the question intent, the item description information associated with the aggregation result, and the image of the item to be identified, and the identification result is output; the identification result also includes the item information related to the question intent.

11. An interactive device for object recognition, comprising: The display module is used to display the target video of the object to be identified, which is input through the interactive interface. as well as The display module is used to use an intelligent agent to identify a target video of an object to be identified and to display the first identification result on the interactive interface; wherein, the first identification result includes feature summary information of the object to be identified.

12. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. Wherein, when one or more programs are executed by one or more processors, the one or more processors implement the method of any one of claims 1 to 10.

13. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 10.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 10.