Object recognition method and apparatus, and device and storage medium
By combining image and text information in media content, candidate object regions are identified and multimodal feature matching is performed, solving the problem of low object recognition accuracy in existing technologies and achieving more efficient object recognition results.
Patent Information
- Application Number
- PCT/CN2025/075307
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-01-26
- Publication Date
- 2025-12-04
AI Technical Summary
Existing technologies struggle to effectively improve the accuracy of object recognition in media content, especially in the area of multimodal information fusion.
Candidate object regions are determined by image information based on media content, and target object regions are determined by combining text information associated with the media content. Matching is performed using text features and visual features, and object recognition is achieved by using a multimodal information fusion method.
It improves the accuracy of object recognition, especially when processing multimodal information, and enhances the ability to identify target objects.
Smart Images

Figure CN2025075307_04122025_PF_FP_ABST
Abstract
Description
Object recognition methods, apparatus, devices and storage media
[0001] This application claims priority to Chinese Patent Application No. 202410696398.0, filed on May 31, 2024, entitled "Object Recognition Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to object identification methods, apparatus, devices, and computer-readable storage media. Background Technology
[0003] With the rapid development of intelligent technology, various forms of media content can greatly enrich people's daily lives. Media content can include objects, such as people and things. How to identify the objects included in media content is a key issue of concern. Summary of the Invention
[0004] In a first aspect of this disclosure, an object recognition method is provided. The method includes: determining a set of first candidate object regions based on image information of media content; determining a target object region from the set of first candidate object regions based on text information associated with the media content; and determining a target object matching the target object region based on text features and visual features of the target object region, wherein the text features are determined based on the text information.
[0005] In a second aspect of this disclosure, an apparatus for object recognition is provided. The apparatus includes: a first determining module configured to determine a set of first candidate object regions based on image information of media content; a second determining module configured to determine a target object region from the set of first candidate object regions based on text information associated with the media content; and a third determining module configured to determine a target object matching the target object region based on text features and visual features of the target object region, wherein the text features are determined based on the text information.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] Figure 2 illustrates a flowchart of an object identification process according to some embodiments of the present disclosure;
[0013] Figure 3 illustrates an example flowchart of object recognition according to some embodiments of the present disclosure;
[0014] Figure 4 shows an example flowchart of object recognition according to some embodiments of the present disclosure;
[0015] Figure 5 shows a schematic diagram of a model structure according to some embodiments of the present disclosure;
[0016] Figure 6 shows a schematic structural block diagram of an apparatus for object recognition according to certain embodiments of the present disclosure;
[0017] Figure 7 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0020] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0021] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0022] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0023] Embodiments of this disclosure propose an object recognition scheme. According to the scheme, a first set of candidate object regions is determined based on image information of media content; a target object region is determined from the first set of candidate object regions based on text information associated with the media content; and a target object matching the target object region is determined based on text features and visual features of the target object region, wherein the text features are determined based on the text information.
[0024] Based on this approach, embodiments of this disclosure can identify target objects in media content using multimodal information such as image information and text information associated with the media content, effectively improving the accuracy of object recognition.
[0025] Example Environment
[0026] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.
[0027] In this example environment 100, electronic device 110 can run an application 120 that supports user interface interaction. Application 120 can be any suitable type of application for user interface interaction. User 140 can view media content based on application 120, where the media content can be short videos, live videos, text and images, and any suitable form of media content. Target user 140 can interact with application 120 via electronic device 110 and / or its attached devices.
[0028] In environment 100 of Figure 1, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.
[0029] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of interface for the target user (such as "wearable" circuitry).
[0030] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support user interface interaction in electronic devices 110.
[0031] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0033] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0034] Example process
[0035] Figure 2 shows a flowchart of an object identification process 200 according to some embodiments of the present disclosure. Process 200 can be implemented at electronic device 110. Process 200 is described below with reference to Figure 1.
[0036] In box 210, electronic device 110 determines a set of candidate object regions based on image information of the media content.
[0037] In some embodiments, the media content can be any suitable type of content such as short videos or live videos, and the media content includes multiple images. The multiple images can include objects, which can be people, things, or any suitable type of object. As an example, the object can be a product.
[0038] In some embodiments, the electronic device 110 may determine a set of candidate object regions based on image information corresponding to all images in the media content.
[0039] In other embodiments, the electronic device 110 can also determine a subset of keyframes from all images in the media content, and determine a set of candidate object regions based on the image information corresponding to these keyframes. Taking Figure 3 as an example, the media content can be video 301. The electronic device can extract keyframes from the images included in this video 301 to obtain a video frame sequence, which is also a subset of keyframes. The electronic device can perform object detection based on the image information of this video frame sequence to obtain a set of first candidate object regions 305.
[0040] The candidate object region is the area in the image corresponding to the object contained in the media content. This region can be displayed in the image in the form of an object box. The object box can be any appropriate shape, such as a rectangle, an irregular shape, etc., which will not be elaborated here.
[0041] In box 220, electronic device 110 determines the target object region from a set of candidate object regions based on text information associated with media content.
[0042] In some embodiments, the text information associated with the media content may include, but is not limited to, at least one of the following: extracting first text content from the image content of the media content; or extracting second text content from the audio content of the media content; or determining third text content based on the descriptive information of the media content. The descriptive information may be a title for the media content, introductory information about objects contained in the media content, etc., and the introductory information may include the type of object, the name of the object, etc. Taking Figure 3 as an example, the text information associated with the media content may include explanatory text 302 obtained after speech recognition of the speech included in the media content, Optical Character Recognition (OCR) text 303 obtained based on image recognition of the media content, title text 304 obtained based on text recognition of the media content, etc.
[0043] In some embodiments, the electronic device 110 can obtain a target image by stitching together a set of images corresponding to a set of candidate object regions, so that the target image includes global information between adjacent frames or multiple images in this set of images. In some embodiments, the stitching method can be any suitable stitching method, such as stitching multiple images into a target image of N*M specifications, where N and M can be set as needed. The electronic device 110 can obtain the target object region output by the first model by inputting the target image and text information associated with the media content into the first model. Taking Figure 2 as an example, the electronic device 110 can perform multimodal subject determination based on the explanatory text 302, OCR text 303, title text 304, and target image to determine the target object region 306. The multimodal subject determination of this disclosure comprehensively considers the text modality information and the image modality information of the media content, which can effectively improve the accuracy of object recognition.
[0044] The following explanation uses the example of the training process of the first model being executed by electronic device 110 to illustrate the training process of the first model. Of course, the training process of the first model can also be executed by other devices, which will not be elaborated here.
[0045] To obtain a trained first model, the electronic device 110 may acquire a first training set, which may include multiple first training samples. Each first training sample may include a target sample image determined by a sample video and sample text information associated with that sample video. In some embodiments, the target sample image may be obtained by stitching together a set of sample images, and this set of sample images may be obtained by sampling the sample video at a predetermined sampling interval. As an example, the electronic device 110 may downsample the sample video by 1 frame / 2 seconds to obtain this set of sample images. For each sample image in this set of sample images, the sample image corresponds to a first labeled object region labeled with the sample object included in the sample image. In some embodiments, the sample text may be an object title, object category, object name, etc.
[0046] For each first training sample in the first training set, the electronic device 110 can input the first training sample in the first training set into the first model to be trained, and obtain a set of predicted object regions output by the first model, and a first target score corresponding to this set of predicted object regions. The electronic device 110 can retain the predicted object regions whose first target score is greater than a predetermined score threshold, and delete the predicted object regions whose score is less than the predetermined score threshold.
[0047] The electronic device 110 can train the first model to be trained based on the comparison between the first labeled object region and the retained predicted object region. After the predetermined training conditions are met, the electronic device 110 can determine that the training of the first model is complete. The predetermined training conditions may be that the loss function reaches its minimum value, the training time reaches a predetermined time, etc., which will not be elaborated here.
[0048] Since media content may include multiple objects, but some of these objects may not be key objects in the media content, in order to accurately identify key objects in the media content, electronic device 110 can further filter out object regions containing key objects from the target object region. Taking Figure 3 as an example, electronic device 110 can track each object contained in the target object region 306 and determine object tracking tracks, which can determine the number of images corresponding to the same object in each image corresponding to the target object region. The more images, the greater the probability that the object is a key object. In some embodiments, after determining the target object region from a set of first candidate object regions, electronic device 110 can determine whether the number of images corresponding to the same object in each image corresponding to the target object region is greater than a predetermined number threshold. In response to the number of images corresponding to the same object not being greater than the predetermined number threshold, electronic device 110 can delete the target object region corresponding to the images of the same object. For example, if there are two images containing object A and ten objects containing object B in the images corresponding to the target object region, then the electronic device 110 can delete the target object region corresponding to the image containing object A to ensure that the electronic device 110 does not perform object recognition on object A in the future.
[0049] Because some objects in images may be occluded or difficult to identify due to shooting angles, in order to improve the accuracy of object recognition in media content, taking Figure 3 as an example, the electronic device 110 can perform query quality judgment to determine the query object box. That is, after determining the target object region 306 from a set of first candidate object regions, the electronic device 110 can determine whether the quality corresponding to the target object region 306 is greater than a predetermined quality threshold. The electronic device 110 can delete the target object region whose quality corresponding to the target object region 306 is not greater than the predetermined quality threshold in response to the target object region 306, thereby filtering out target object regions that reduce the accuracy of object recognition.
[0050] In box 230, electronic device 110 determines the target object that matches the target object region based on text features and visual features of the target object region, with the text features determined based on text information.
[0051] In some embodiments, in order to filter out noise information in the text information and improve the accuracy of object recognition, the electronic device 110 may determine text features based on the text information before determining the target object matching the target object region based on text features and visual features of the target object region.
[0052] As an example, electronic device 110 can obtain text features output by the third model by inputting text information into the third model, where the text features are structured descriptive features generated based on the text information and with noise information filtered out.
[0053] As another example, electronic device 110 can obtain a set of candidate images associated with text information, wherein the time interval between the first appearance of this set of candidate images in the media content and the second appearance of the text information in the media content is less than a predetermined interval threshold. Electronic device 110 can use this set of candidate images as target prompts, and obtain the text features output by the third model by inputting the text information and the target prompts as input information into the third model.
[0054] Taking Figure 3 as an example, the electronic device can perform one-way object recall based on the visual features 307 of the target object region and the object box feature library, and perform another-way object recall based on the text features 308 and the object text features. Finally, the target object is determined based on the objects recalled by the two channels.
[0055] In some embodiments, when performing one-way object recall based on text features and object text features, the electronic device 110 can determine a first candidate object based on the comparison results between the text features and various features in the text feature library. As an example, the electronic device 110 can determine a first similarity between the text features and various features in the text feature library. The electronic device 110 can determine the object corresponding to the feature in the text feature library whose first similarity is greater than a predetermined similarity threshold as the first candidate object.
[0056] In some embodiments, when performing one-way object retrieval based on the visual features of the target object region and an object bounding box feature library, the electronic device 110 can determine a second candidate object based on the comparison results between the visual features of the target object region and each feature in the feature library corresponding to the object region. As an example, the electronic device 110 can determine a second similarity between the visual features of the target object region and each feature in the feature library corresponding to the object region. The electronic device 110 can determine objects corresponding to features in the feature library corresponding to the object region whose second similarity is greater than a predetermined similarity threshold as second candidate objects.
[0057] In some embodiments, the electronic device 110 may determine a target object that matches the target object region based on a first candidate object and a second candidate object.
[0058] As an example, electronic device 110 can determine whether a first candidate object and a second candidate object are the same, so as to identify the same candidate object as the target object that matches the target object region.
[0059] As another example, when determining the target object based on the objects recalled by two channels, the electronic device 110 can determine a set of fourth candidate objects that match the target object region based on the first candidate object and the second candidate object. In some embodiments, the electronic device 110 can obtain the object features corresponding to this set of fourth candidate objects. Taking Figure 3 as an example, the electronic device 110 can perform video-object relevance ranking on this set of fourth candidate objects based on object features and visual multi-dimensional features, and finally determine the target object 309 based on the ranking result. In some embodiments, the electronic device 110 can determine the similarity of a set of fourth candidate objects by inputting text information, image information, and object features corresponding to a set of fourth candidate objects into a fourth model. The higher the similarity of the fourth candidate objects, the greater the probability that it corresponds to the target object. The electronic device 110 can determine the target object based on the similarity of a set of fourth candidate objects. As an example, the electronic device 110 can determine the fourth candidate object with the highest similarity in this set of fourth candidate objects as the target object. As another example, the electronic device 110 can determine the fourth candidate objects in this set of fourth candidate objects whose similarity is greater than a predetermined similarity threshold as the target object.
[0060] Using Figure 4 as an example, the process of determining each feature in the text feature library and each feature in the feature library corresponding to the object region will be explained below.
[0061] In some embodiments, the electronic device 110 may store various object images in a library so that the electronic device can obtain object images from the object library.
[0062] Electronic device 110 can input object images from an object library and associated text information into a second model to obtain a second candidate object region output by the second model. The text information associated with the object image information can be the object's name, category, or other object description. Taking Figure 4 as an example, the electronic device can input multiple object images, object titles, and object categories into the second model. The electronic device can use the second model to perform object detection 401 on the multiple object images to obtain object regions within the multiple object images. The electronic device can use this object region, along with the object title 402 and object category 403, to perform object localization to obtain the second candidate object region output by the second model. Electronic device 110 can determine the visual features of the second candidate object region. Based on the visual features of the second candidate object region, electronic device 110 can determine the features in the feature library corresponding to the object region, i.e., determine the features in the image feature library.
[0063] To obtain a trained second model, the electronic device 110 can acquire a second training set, which may include multiple second training samples. Each second training sample may include a sample image corresponding to a sample object and sample text corresponding to the sample object. For each sample image, the sample image corresponds to a labeled second labeled object region of the sample object within that image. The sample text may include the object title, object category, object name, etc. For each second training sample, the electronic device 110 can input the second training sample into the second model to be trained, obtaining a set of predicted object regions output by the second model, and a second target score corresponding to this set of predicted object regions. The electronic device 110 can retain predicted object regions with a second target score greater than a predetermined score threshold and delete predicted object regions with a second target score less than the predetermined score threshold.
[0064] The electronic device 110 can train the second model to be trained based on the comparison between the second labeled object region and the retained predicted object region. After the predetermined training conditions are met, the training of the second model is completed. The predetermined training conditions may be that the loss function reaches its minimum value, the training time reaches a predetermined time, etc., which will not be elaborated here.
[0065] Using Figure 5 as an example, the second model may include a text encoder 501, an image editor 502, and a modal interactive decoder 503. The text encoder 501 is used to obtain the text representation corresponding to the sample text, and the image editor 502 is used to obtain the image representation corresponding to the sample image. The modal interactive decoder 503 is used to comprehensively predict the set of predicted object regions corresponding to the sample image input into the second model based on this text representation and image representation.
[0066] In some embodiments, the electronic device 110 can determine the text features of a third candidate object included in the object image based on text information associated with the image information of the object image. In some embodiments, the electronic device 110 can obtain the text features of the third candidate object included in the object image by inputting the text information associated with the image information of the object image into a third model, wherein the text features of the third candidate object are structured descriptive features generated from the text information associated with the image information of the object image after filtering out noise. The electronic device 110 can determine various features in the text feature library based on the text features of the third candidate object, i.e., determine the features in the text feature library.
[0067] The process of obtaining the visual features of the target object region is explained below.
[0068] In some embodiments, before determining the target object matching the target object region based on text features and visual features of the target object region, the electronic device 110 can obtain the visual features of the target object region output by inputting the image corresponding to the target object region into a visual feature model. In some embodiments, the visual feature model can employ a convolutional network or an attention mechanism converter to extract the visual features of the image corresponding to the target object region input into the visual feature model.
[0069] In some embodiments, to obtain a trained visual feature model, the electronic device 110 can obtain a training set and train the visual feature model based on the training set. In some embodiments, the electronic device 110 can determine a set of first sample images from a sample video and a second sample image from an object library, wherein the first and second sample images contain sample objects. The sample video can be a video captured in a real scene. The images in the object library can be images containing objects in non-real scenes, such as rendered or generated images. The electronic device 110 can use this set of first and second sample images as a first candidate sample image pair, wherein each image in the first candidate sample image pair is labeled with an object region. The electronic device 110 can determine a first similarity between a first feature corresponding to the object region in the first sample image and a second feature corresponding to the second sample image. To ensure that the object has a consistent representation in both real and non-real scenes, the electronic device 110 can add the first candidate sample image pair with a first similarity less than a first threshold but greater than a second threshold to the training set.
[0070] In some embodiments, after determining a set of first sample images from a sample video, the electronic device 110 may use the set of first sample images and a feedback image as a second candidate sample image pair, wherein each image in the second candidate sample image pair is labeled with an object region. The feedback image is an image associated with an object included in the sample video, obtained for that sample video. The feedback image may be an image related to an object commented on by a user while watching the sample video. For example, the electronic device 110 may receive other images including object B sent by user A in the comment section of a sample video containing object B. The electronic device 110 may determine a third feature corresponding to the object region in the feedback image, and determine a second similarity between the first feature and the third feature corresponding to the object region in the set of first sample images. The electronic device 110 may add candidate sample image pairs with a second similarity less than a first threshold but greater than a second threshold to the training set.
[0071] After obtaining the training set, the electronic device 110 can train the visual feature model to be trained based on this training set. In some embodiments, the visual feature model may include a cross-domain alignment module, which is used to align the video domain with the object domain, also known as the alignment of the same object in real and non-real scenes. In some embodiments, the electronic device 110 can determine a triplet contrastive loss function and train the visual feature model to be trained based on the triplet contrastive loss function, wherein the triplet may include an anchor point, a positive sample point, and a negative sample point. In this embodiment of the present disclosure, the anchor point can be an image in a non-real scene, i.e., an image in an object library, and the positive and negative sample points are images in a real scene, i.e., images in a sample video. During the training of the visual feature model, the feature distance between the anchor point and the positive sample can be shortened, and the feature distance between the anchor point and the negative sample can be lengthened. As an example, the electronic device 110 can determine that the training of the visual feature model to be trained is complete when the triplet contrastive loss function is determined to be at its minimum value.
[0072] Based on this approach, embodiments of this disclosure can identify target objects in media content using multimodal information such as image information and text information associated with the media content, effectively improving the accuracy of object recognition.
[0073] Example devices and equipment
[0074] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. FIG6 shows a schematic structural block diagram of an apparatus 600 for providing media content according to certain embodiments of this disclosure. Apparatus 600 may be implemented as or included in the electronic device 110 discussed above. The various modules / components in apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0075] As shown in Figure 5, the device 600 includes a first determining module 610 configured to determine a set of first candidate object regions based on image information of media content; a second determining module 620 configured to determine a target object region from the set of first candidate object regions based on text information associated with the media content; and a third determining module 630 configured to determine a target object matching the target object region based on text features and visual features of the target object region, wherein the text features are determined based on text information.
[0076] In some embodiments, the second determining module 620 is specifically configured to obtain a target image by stitching together a set of images corresponding to a set of first candidate object regions; and to obtain the target object region output by the first model by inputting the target image and text information associated with media content into the first model.
[0077] In some embodiments, the apparatus 600 further includes a fourth determining module, configured to: determine whether the number of images corresponding to the same object in each image corresponding to the target object region is greater than a predetermined number threshold; and a first deleting module, configured to: delete the target object region corresponding to the image of the same object in response to the fact that the number of images corresponding to the same object is not greater than the predetermined number threshold.
[0078] In some embodiments, the apparatus 600 further includes a fifth determining module configured to: determine whether the quality corresponding to the target object region is greater than a predetermined quality threshold; and a second deleting module configured to: delete the target object region in response to the target object region's quality not being greater than the predetermined quality threshold.
[0079] In some embodiments, the third determining module 630 is specifically configured to: determine a first candidate object based on the comparison results of text features and various features in the text feature library; determine a second candidate object based on the comparison results of visual features of the target object region and various features in the feature library corresponding to the object region; and determine a target object that matches the target object region based on the first candidate object and the second candidate object.
[0080] In some embodiments, the apparatus 600 further includes a first obtaining module configured to: obtain an object image from an object library; a second obtaining module configured to: obtain a second candidate object region output by the second model by inputting the object image from the object library and text information associated with the image information of the object image into the second model; a sixth determining module configured to: determine the visual features of the second candidate object region; and a seventh determining module configured to: determine the features in the feature library corresponding to the object region based on the visual features of the second candidate object region.
[0081] In some embodiments, the apparatus 600 further includes an eighth determining module configured to: determine text features of a third candidate object included in the object image based on text information associated with image information of the object image; and a ninth determining module configured to: determine each feature in the text feature library based on the text features of the third candidate object.
[0082] In some embodiments, the apparatus 600 further includes a third obtaining module configured to obtain text features output by the third model by inputting text information into the third model, wherein the text features are structured descriptive features generated based on the text information.
[0083] In some embodiments, the third obtaining module is specifically configured to obtain a set of candidate images associated with text information, wherein the time interval between the first time when the set of candidate images appears in the media content and the second time when the text information appears in the media content is less than a predetermined interval threshold; to use the set of candidate images as target prompts; and to obtain the text features output by the third model by inputting the text information and the target prompts as input information into the third model.
[0084] In some embodiments, the third determining module is specifically configured to: determine a set of fourth candidate objects matching the target object region based on text features and visual features of the target object region; obtain object features corresponding to the set of fourth candidate objects; determine the similarity corresponding to the set of fourth candidate objects by inputting text information, image information, and object features corresponding to the set of fourth candidate objects into a fourth model; and determine the target object based on the similarity of the set of fourth candidate objects.
[0085] In some embodiments, the text information includes at least one of the following: extracting first text content from the image content of the media content; or extracting second text content from the audio content of the media content; or determining third text content based on the descriptive information of the media content.
[0086] In some embodiments, the apparatus 600 further includes a fourth obtaining module, configured to: obtain visual features of the target object region output by the visual feature model by inputting an image corresponding to the target object region into the visual feature model.
[0087] In some embodiments, the apparatus 600 further includes a tenth determining module configured to: determine a set of first sample images containing sample objects from the sample video; an eleventh determining module configured to: determine a second sample image containing sample objects from an object library; use the set of first sample images and the second sample images as a first candidate sample image pair, wherein each image in the first candidate sample image pair is labeled with an object region; the eleventh determining module configured to: determine a first similarity between a first feature corresponding to the object region in the set of first sample images and a second feature corresponding to the second sample image; and an adding module configured to: add the first candidate sample image pair with a first similarity less than a first threshold but greater than a second threshold to the training set.
[0088] In some embodiments, the apparatus 600 further includes a fifth obtaining module configured to: obtain a feedback image for a sample video, the feedback image being an image associated with an object included in the sample video; a twelfth determining module configured to: use a set of first sample images and the feedback image as a second candidate sample image pair, wherein each image in the second candidate sample image pair is labeled with an object region; a thirteenth determining module configured to: determine a third feature corresponding to the object region in the feedback image; a fourteenth determining module configured to: determine a second similarity between the first feature and the third feature corresponding to the object region in the set of first sample images; and a second adding module configured to: add candidate sample image pairs with a second similarity less than a first threshold but greater than a second threshold to the training set.
[0089] The units included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 600 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0090] Figure 7 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 700 shown in Figure 7 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 700 shown in Figure 7 can be used to implement the electronic device 110 shown in Figure 1.
[0091] As shown in Figure 7, the electronic device 700 is in the form of a general-purpose electronic device. Components of the electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage devices 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.
[0092] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 700.
[0093] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0094] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0095] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0096] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0097] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0098] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0099] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0101] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An object recognition method, comprising: Based on the image information of the media content, a set of first candidate object regions is determined; Based on the text information associated with the media content, a target object region is determined from the set of first candidate object regions; as well as Based on text features and visual features of the target object region, a target object matching the target object region is determined, wherein the text features are determined based on the text information.
2. The method of claim 1, wherein determining the target object region from the set of first candidate object regions based on text information associated with the media content comprises: The target image is obtained by stitching together a set of images corresponding to the first set of candidate object regions; as well as By inputting the target image and text information associated with the media content into the first model, the target object region output by the first model is obtained.
3. The method according to claim 1, after determining the target object region from the set of first candidate object regions, the method further includes: Determine whether the number of images corresponding to the same object in each image corresponding to the target object region is greater than a predetermined number threshold; In response to the fact that the number of images corresponding to the same object is not greater than the predetermined number threshold, the target object region corresponding to the images of the same object is deleted.
4. The method according to claim 1, after determining the target object region from the set of first candidate object regions, the method further includes: Determine whether the quality corresponding to the target object region is greater than a predetermined quality threshold. as well as In response to the fact that the quality of the target object region is not greater than the predetermined quality threshold, the target object region is deleted.
5. The method according to claim 1, wherein determining the target object matching the target object region based on text features and visual features of the target object region comprises: Based on the comparison results between the text features and the features in the text feature library, the first candidate object is determined; Based on the comparison results between the visual features of the target object region and each feature in the feature library corresponding to the object region, a second candidate object is determined; Based on the first candidate object and the second candidate object, a target object matching the target object region is determined.
6. The method according to claim 5, wherein each feature in the feature library corresponding to the object region is determined based on the following process: Obtain object images from the object library; By inputting object images from the object library and text information associated with the image information of the object images into the second model, the second candidate object region output by the second model is obtained; Determine the visual features of the second candidate object region; as well as Based on the visual features of the second candidate object region, the features in the feature library corresponding to the object region are determined.
7. The method of claim 6, wherein each feature in the text feature library is determined based on the following process: Based on the text information associated with the image information of the object image, determine the text features of the third candidate object included in the object image; Based on the text features of the third candidate object, each feature in the text feature library is determined.
8. The method according to claim 1, wherein before determining a target object matching the target object region based on text features and visual features of the target object region, the method comprises: By inputting the text information into a third model, the text features output by the third model are obtained, wherein the text features are structured descriptive features generated based on the text information.
9. The method according to claim 8, wherein obtaining the text features output by the third model by inputting the text information into the third model comprises: Obtain a set of candidate images associated with the text information, wherein the time interval between the first time the set of candidate images appear in the media content and the second time the text information appears in the media content is less than a predetermined interval threshold; Use the aforementioned set of candidate images as target prompts; as well as By inputting the text information and the target prompt as input information into the third model, the text features output by the third model are obtained.
10. The method of claim 1, wherein determining the target object matching the target object region based on text features and visual features of the target object region comprises: Based on the text features and the visual features of the target object region, a set of fourth candidate objects matching the target object region is determined; Obtain the object features corresponding to the fourth set of candidate objects; The similarity between the set of fourth candidate objects is determined by inputting the text information, the image information, and the object features corresponding to the set of fourth candidate objects into the fourth model. as well as The target object is determined based on the similarity of the fourth set of candidate objects.
11. The method according to claim 1, wherein the text information includes at least one of the following: Extract the first text content from the image content of the media content; or Extracting second text content from the audio content of media content; or The content of the third text is determined based on the descriptive information of the media content.
12. The method according to claim 1, further comprising, before determining the target object matching the target object region based on text features and visual features of the target object region: By inputting the image corresponding to the target object region into the visual feature model, the visual features of the target object region output by the visual feature model are obtained.
13. The method of claim 12, wherein the training set for training the visual feature model is determined in the following manner: From the sample video, a set of first sample images is determined, wherein the first sample images contain sample objects; Determine a second sample image containing the sample object from the object library; The first set of sample images and the second set of sample images are used as a first candidate sample image pair, wherein each image in the first candidate sample image pair is marked with an object region; Determine the first similarity between the first feature corresponding to the object region in the first set of first sample images and the second feature corresponding to the second sample image; and The first candidate sample image pair with a first similarity less than a first threshold but greater than a second threshold is added to the training set.
14. The method of claim 13, further comprising, after determining a first set of sample images from the sample video: The first set of sample images and the feedback image are used as a second candidate sample image pair, wherein each image in the second candidate sample image pair is marked with an object region, and the feedback image is an image obtained for the sample video that is associated with the object contained in the sample video; Determine the third feature corresponding to the object region in the feedback image; Determine the first feature corresponding to the object region in the set of first sample images and the second similarity of the third feature; as well as Candidate sample image pairs whose second similarity is less than the first threshold but greater than the second threshold are added to the training set.
15. An apparatus for generating media content, comprising: The first determining module is configured to determine a set of first candidate object regions based on image information of media content; The second determining module is configured to determine a target object region from the set of first candidate object regions based on text information associated with the media content; as well as The third determining module is configured to determine a target object that matches the target object region based on text features and visual features of the target object region, wherein the text features are determined based on the text information.
16. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processor.
17. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 14.
18. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Image text processing method and device, readable medium and electronic equipment
CN115331228A
Object information identification method, related device, equipment and storage medium
CN116563905A
Information retrieval method and device, equipment, program product and storage medium
CN116975340A
Media information identification method and device, computer equipment and storage medium
CN117725234A
Methods and systems to identify an object in content
US20190080175A1