Image description method and device, electronic device, and storage medium
By acquiring information about the subject, details, and background of an image, multi-dimensional descriptive information is generated, solving the problem of poor descriptive accuracy in existing technologies and improving the comprehensiveness and relevance of image descriptions.
Patent Information
- Application Number
- CN202211174872.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Existing image description methods mainly focus on subject information while ignoring background and detail information, resulting in poor description accuracy. They also cannot freely select the target subject, leading to unfocused and poorly targeted description information.
By acquiring the subject information, detail information, and background information of an image, and combining multiple dimensions to generate descriptive information, users can freely select the target subject for description, generating highly targeted descriptive information.
It improves the comprehensiveness, accuracy, and relevance of image descriptions, enhances the matching and authenticity between descriptive information and image content, and solves the problem of unfocused descriptive information.
Smart Images

Figure CN115512213B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and more specifically, to an image description method and apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0002] During image processing, descriptive information can be generated to enable users to read and process the image, thereby improving the user experience.
[0003] In related technologies, the content of an image can be described based on the objects identified in the image. However, this method can only describe the image based on the subject information, which has certain limitations and relatively poor accuracy.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this disclosure is to provide an image description method, apparatus, electronic device, and storage medium, thereby overcoming, to at least a certain extent, the problem of poor accuracy of image description information caused by the limitations and defects of related technologies.
[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0007] According to a first aspect of this disclosure, an image description method is provided, comprising: acquiring an image to be processed, performing target detection on the image to be processed, and determining subject information of at least one subject contained in the image to be processed; acquiring background information based on the remaining image in the image to be processed excluding each of the subjects; performing classification detection on the subjects in the image to be processed, and determining detailed information of each of the subjects; determining a target subject based on each of the subjects, and generating description information of the image to be processed based on the subject information, the detailed information, and the background information of the target subject.
[0008] According to a second aspect of this disclosure, an image description method is provided, comprising: acquiring an image to be processed and generating description information of the image to be processed; the description information being generated based on subject information, detail information, and background information of a target subject in the image to be processed; and playing the description information to read the image to be processed.
[0009] According to a third aspect of the present disclosure, an image description apparatus is provided, comprising: a subject information acquisition module configured to acquire a to-be-processed image, and perform target detection on the to-be-processed image to determine subject information of at least one subject contained in the to-be-processed image; a background information acquisition module configured to acquire background information according to remaining images in the to-be-processed image except for each of the subjects; a detail information determination module configured to perform classification detection on the subjects in the to-be-processed image to determine detail information of each of the subjects; and a description information generation module configured to determine a target subject based on each of the subjects, and generate description information of the to-be-processed image according to the subject information, the detail information and the background information of the target subject.
[0010] According to a fourth aspect of the present disclosure, an image description apparatus is provided, comprising: a description information generation module configured to acquire a to-be-processed image, and generate description information of the to-be-processed image; the description information is generated according to subject information, detail information and background information of a target subject in the to-be-processed image; and an information playing module configured to play the description information to read the to-be-processed image.
[0011] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the image description method of the first aspect and the second aspect and possible implementation manners thereof by executing the executable instructions.
[0012] According to a sixth aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the image description method of the first aspect and the second aspect and possible implementation manners thereof.
[0013] In the technical solutions provided in the embodiments of the present disclosure, on one hand, the description information for describing the content of the to-be-processed image is generated from multiple dimensions by combining the subject information, the detail information and the background information of the to-be-processed image, compared with the prior art, the limitation that the description information can only be generated according to the subject information is avoided, the comprehensiveness is increased by describing from multiple dimensions, the description range is increased and the accuracy of the description information is improved. On the other hand, since the description information for the target subject can be generated by combining multiple dimensions in the to-be-processed image, the matching and rationality of the description information and the content of the to-be-processed image are improved, the authenticity is increased, the effect and accuracy of the image description are improved, and the application range is increased. On the other hand, since the target subject can be freely selected to generate the description information for the target subject, the situation that the description information is not focused in the related art is avoided, and the pertinence and accuracy of the description information are improved.
[0014] It should be understood that the foregoing general description and the following detailed description are only examples and explanatory, and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure. It is readily apparent to one skilled in the art that the following description is merely exemplary and explanatory, and does not limit the present disclosure, and other embodiments can be readily derived from the drawings without departing from the scope of the present disclosure.
[0016] Figure 1 A schematic diagram showing an application scenario of an image description method to which embodiments of the present disclosure can be applied.
[0017] Figure 2 A schematic diagram schematically showing an image description method in embodiments of the present disclosure.
[0018] Figure 3 A schematic diagram schematically showing determination of subject information in embodiments of the present disclosure.
[0019] Figure 4 A schematic diagram schematically showing acquisition of background information in embodiments of the present disclosure.
[0020] Figure 5 A schematic diagram schematically showing a flow of acquisition of detail information in embodiments of the present disclosure.
[0021] Figure 6 A schematic diagram schematically showing determination of a target subject and a reference subject in embodiments of the present disclosure.
[0022] Figure 7 A schematic diagram schematically showing a flow of generation of description information in embodiments of the present disclosure.
[0023] Figure 8 A schematic diagram schematically showing determination of a target action in embodiments of the present disclosure.
[0024] Figure 9 A schematic diagram schematically showing a flow of another image description method in embodiments of the present disclosure.
[0025] Figure 10 A schematic diagram schematically showing a specific flow of reading of an image in embodiments of the present disclosure.
[0026] Figure 11 A block diagram schematically showing an image description apparatus in embodiments of the present disclosure.
[0027] Figure 12 A block diagram schematically showing another image description apparatus in embodiments of the present disclosure.
[0028] Figure 13 A block diagram of an electronic device in embodiments of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0029] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations. In the following description, numerous specific details are provided to give a thorough understanding of implementations of the present disclosure. One skilled in the relevant art will recognize, however, that the implementations of the present disclosure can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures have not been described in detail so as not to obscure the aspects of the present disclosure.
[0030] Furthermore, the accompanying drawings are only schematic and are non-limiting detailed representations of implementations. Like references numerals can be used to denote like parts throughout the figures and description thereof. Some of the figures can be schematic or schematic-functional. Some of the figures can be block diagrams, functional block diagrams, flow diagrams, illustrations of hardware or software architectures, user interfaces, user interface screens, user interface displays, user interface displays, or combinations thereof. In some cases, details have not been shown in order to not obscure the examples described herein. The term "coupled" is used herein to express either a direct or indirect connection between entities. The term "coupled" can include have physical or electrical association or connection, or can include logical association or connection.
[0031] In the related art, the class of an object in an image can be accurately identified, such as identifying a person, a horse, and the like. Further, the subject is organized into a sentence through certain language rules, and the subject detail description (the main details such as the subject position relationship, the background, the expression, the action, and the like are easily ignored). For example, the image is identified as a person riding a horse, in which only the subject is described, but the background, the expression, and even the action, the color, and the like are ignored. For another example, the image is identified as a person sitting in a car, which is easy to lose background information, discard other subject information, and the like. The main problems in the related art are as follows: mainly focusing on the target subject, easily ignoring the background information and the detail information; and unable to select the subject, thus being less targeted.
[0032] To solve the technical problems in the related art, an image description method is provided in embodiments of the present disclosure, which can be applied to any scene in which the image cannot be viewed and needs to be played and read. Figure 1 A schematic diagram of a system architecture to which the image description method and device of embodiments of the present disclosure can be applied is shown.
[0033] like Figure 1 As shown, terminal 101 can be a smart device with image processing capabilities, such as a smartphone, computer, tablet, smart speaker, smartwatch, in-vehicle device, wearable device, monitoring device, etc. The terminal may include a display module for displaying the image. The image to be processed can be a captured image, each frame of a captured video, a received image, or a stored image; there is no limitation here.
[0034] In this embodiment, terminal 101 may include memory 102 and processor 103. The memory stores images, and the processor processes the images, such as performing subject recognition, background recognition, and acquiring detailed information. Memory 102 may store an image 104 to be processed. Terminal 101 retrieves the image 104 to be processed from memory 102 and sends it to processor 103. Processor 103 performs target detection on the image to be processed to determine subject information; acquires background information based on the remaining image in the image to be processed excluding the subjects; performs classification detection on the subjects in the image to be processed to determine detailed information of each subject; determines a target subject based on each subject; and generates description information 105 of the target subject based on the subject information, detailed information, and background information, so that the description information can be played.
[0035] It should be noted that the image description method provided in this embodiment can be executed by terminal 101. The image description device can also be located in the terminal.
[0036] Figure 2 The image description method in this embodiment is illustrated schematically and can be applied to any type of application scenario where images cannot be viewed. For example, it can be applied to driving scenarios, sports scenarios, and scenarios where the content of an image cannot be viewed due to various factors, such as early childhood education scenarios, elderly viewing scenarios, visually impaired scenarios, etc. In addition, it can also be applied to application scenarios where images cannot be displayed in real time due to network problems, etc. (See reference...) Figure 2 As shown, the specific steps include:
[0037] Step S210: Obtain the image to be processed, and perform target detection on the image to be processed to determine the subject information of at least one subject contained in the image to be processed;
[0038] Step S220: Obtain background information from the remaining image in the image to be processed, excluding each of the main subjects.
[0039] Step S230: Classify and detect the main subjects in the image to be processed to determine the detailed information of each subject;
[0040] Step S240: Determine the target subject based on each of the subjects, and generate description information of the image to be processed according to the subject information, detail information and background information of the target subject.
[0041] In this embodiment of the disclosure, the image to be processed can be any type of image. For example, it can be a static image, a dynamic image, etc.
[0042] To accurately describe an image, subject recognition can be performed to determine the subject information it contains. Subject information can include subject category and subject location information. Subject categories can include, but are not limited to, people, animals, sky, grass / vegetation, buildings, vehicles, etc. Location information refers to the coordinates of each subject in the image, such as the coordinates of the top-left corner, bottom-right corner, and center position.
[0043] Furthermore, the main subject in the image to be processed can be removed, and the remaining image excluding the subject can be padded to obtain a padded image. Based on the padded image, the background category can be identified to determine the background information. The background information can be, for example, grass, beach, etc.
[0044] After obtaining the background information, the identified subject in the image is subjected to detailed analysis to obtain its detailed information. This detailed information can include positional relationships and attribute information. Attribute information can include actions and other attributes. Actions can be individual actions of the subject itself or related actions to other subjects, such as standing or hugging another subject. Other attribute information can include facial expressions, colors, etc. If the subject is not classified as a human, it is not necessary to obtain facial expressions.
[0045] Since an image to be processed can contain multiple subjects, to accurately describe its content, one subject can be selected as the target subject. Based on the target subject's subject information, detailed information, and background information, descriptive information is generated to describe the image's content. This descriptive information represents the semantic information of the image, and may include the target subject's subject information, detailed information, and background information, as well as detailed information from other reference subjects. Therefore, to avoid inconvenience for users viewing the image, the descriptive information can be played back, enabling barrier-free reading of the image.
[0046] Figure 2The technical solution provided here, on the one hand, generates descriptive information from multiple dimensions to describe the content of the image to be processed by combining the subject information, detail information, and background information of the image to be processed, thereby achieving image reading. Compared with existing technologies, this avoids the limitation of generating descriptive information only based on subject information, and the description from multiple dimensions increases comprehensiveness and improves the accuracy of the descriptive information. On the other hand, because it can combine multiple dimensions of the image to be processed to generate descriptive information targeting the target subject, it improves the matching and rationality of the descriptive information with the content of the image to be processed, increases realism, improves the effect and accuracy of image description, and expands the application scope. Furthermore, because the target subject can be freely selected to generate descriptive information targeting the target subject, it avoids the lack of focus in the descriptive information in related technologies, and improves the targeting and accuracy of the descriptive information.
[0047] Next, refer to Figure 2 As shown, the specific steps of the image description method in the embodiments of this disclosure will be described in detail.
[0048] In step S210, an image to be processed is acquired, and target detection is performed on the image to be processed to determine the subject information of at least one subject contained in the image to be processed.
[0049] In this embodiment, the image to be processed can be an image captured by the camera module of the terminal, a frame from a video recording, or an image sent by another user in an instant messaging application; no limitation is made here. The terminal can be any of a smartphone, digital camera, smartwatch, wearable device, or in-vehicle device, as long as it can perform image storage, image processing, and image display. A smartphone is used as an example here. The image to be processed can be of various types, such as dynamic or static images.
[0050] The image to be processed may contain at least one subject, which represents the content of the image. The subject can be any type of object within the image, such as a person, animal, building, or other type of object, depending on the specific requirements. The subject can be located anywhere within the image; there are no restrictions on its position.
[0051] After acquiring the image to be processed, object detection can be performed. Object detection can be understood as the sum of object recognition and object localization, used to identify all objects of interest in the image and determine their category and location. For example, an object detection model can be used to extract features from the image to be processed, and the feature vectors can be fitted and classified to obtain subject information for each subject in the image. Subject information may include the subject category and location information of the subject in the image.
[0052] The object detection model can be trained on sample images. Specifically, the location and category of the objects to be detected in the sample images can be pre-labeled to obtain the labeled location and category of the objects. The parameters of the object detection model are updated using the training sample set and the labeled location and category of the objects to train the object detection model, resulting in a trained object detection model. The subject category can be represented by the subject name, and the location information can be coordinate information, i.e., multiple coordinates of the subject bounding box, such as, but not limited to, the coordinates of the top left corner, the bottom right corner, and the center.
[0053] For example, when performing object detection on an image to be processed, the resulting subject information can include the subject category and location information, as shown in Table 1:
[0054] Table 1
[0055]
[0056] As shown in Table 1, the image to be processed can contain multiple subjects, such as subject 1, subject 2, and subject n, and the subject information of these subjects can be different. Specifically, the subject categories can be the same or different; for example, subject 1 is a person, and subject 2 and subject n are animals. The position information of the multiple subjects can also be different, which is not specifically limited here.
[0057] refer to Figure 3 As shown, the image to be processed 301 can be input into the target detection model 304, and the output will be the subject category 3021 (person) and location information 3022 corresponding to the subject 302 in the image to be processed 301, and the subject category 3031 (box) and location information 3032 corresponding to the subject 303.
[0058] In step S220, background information is obtained from the remaining images in the image to be processed, excluding the subjects.
[0059] In this embodiment of the disclosure, after obtaining all subjects in the image to be processed, the remaining images (excluding the subjects) can be identified as the remaining images, and background information of the images to be processed can be obtained based on the remaining images. Here, background information can be a background category.
[0060] For example, the remaining image can be identified to obtain the background category. Figure 4 The flowchart for determining background information is illustrated in the diagram. (Refer to...) Figure 4 As shown, the main steps include:
[0061] In step S410, each of the main subjects is extracted to obtain the remaining image, and the background of the remaining image is filled in to obtain the filled image;
[0062] In step S420, the completed image is identified using a target detection algorithm to obtain the background category as background information.
[0063] In some embodiments, the remaining image can be input into a neural network model to extract features, obtaining feature information of the remaining image. This feature information includes pixel location information and image features. The feature information is then input into the neural network model to identify missing regions (i.e., the regions containing the main subject). The identified missing regions and feature information are input into the neural network model, and the semantics corresponding to the missing content are determined based on the feature information of each pixel. The semantics corresponding to the missing content are then input into the neural network model, combined with the feature information extracted by the feature extraction layer, to complete the image, resulting in the completed remaining image. Alternatively, other methods can be used to obtain the completed image; specific limitations are not specified here.
[0064] Furthermore, the padded image can be identified using an object detection model to determine its category, thereby identifying background information based on the background category. For example, the padded image's features can be extracted using an object detection model, and the feature vectors can be classified and fitted to determine the padded image's category, with the background category being identified as the background information. For instance, the background category can be identified as grass, water, sand, or other categories. In this embodiment of the disclosure, by obtaining the background information of the image to be processed, the category of the scene in which the subject in the image is located can be accurately determined, facilitating an accurate description of the image's characteristics in conjunction with the scene category. (Reference) Figure 6 As shown, background information 606 can be grassland.
[0065] Next, in step S230, the subjects in the image to be processed are classified and detected to determine the detailed information of each subject.
[0066] In this embodiment, each subject identified in step S210 can be further classified and detected, and detailed analysis can be performed on each subject to obtain detailed information about each subject. The detailed information may include positional relationship information and attribute information of the subjects. Positional relationship information is used to identify the positional relationship between different subjects in the same image and can be used to describe the correlation between different subjects. For example, it can be the positional relationship information between each subject in the image to be processed and any other subject. Positional relationship information may include the positional relationship (i.e., orientation) and distance between subjects. Attribute information can be used to describe the state attributes of each subject, specifically including but not limited to action, color, expression, and other attribute information. Other attribute information may include but not limited to size, specifications, volume, etc., specifically determined according to actual needs.
[0067] In some embodiments, detailed information of the subject can be determined individually based on the image to be processed. Specifically, the subject in the image to be processed can be classified and detected to obtain detailed information. In addition, to improve the accuracy of the subject's detailed information, association analysis can be performed based on the context image of the image to be processed as an auxiliary judgment to accurately obtain the subject's detailed information, thereby accurately describing the background and motion changes of the image to be processed. For example, for dynamic images, there is a context image, so the context image can be used to assist in analyzing the detailed information of each subject in the image to be processed. First, the process of recognizing the subject in the image to be processed by combining the context image will be explained.
[0068] The context image can be an image associated with the image to be processed. The content of the context image can be similar to or the same as the image to be processed, or it can be different. The context image can contain preceding and following images, where the preceding image is an image that occurs before the image to be processed, and the following image is an image that occurs after the image to be processed. The context image can be adjacent to or not adjacent to the image to be processed, as long as a time threshold is met. Furthermore, preceding and following images can coexist, or only one of them can exist, depending on the specific circumstances.
[0069] In some embodiments, it can first be determined whether the image to be processed has a context image. This can be determined based on whether the image to be processed has an associated image and the generation time of all images, or it can be determined based on the type of the image to be processed. For example, when the image to be processed is a moving image, a context image can be considered to exist. Upon determining the existence of a context image, it can be determined whether the time difference between the context image and the image to be processed is less than a time threshold. If it is less than the time threshold, then the context image is used to assist in the analysis of the detailed information of the subject in the image to be processed. The time threshold can be determined according to actual needs.
[0070] After determining the context image, it can be analyzed to obtain its subject information and background information. Specific steps include: performing subject recognition on the context image to obtain the subject information of the subject in the context image; and performing object detection on the remaining image in the context image (excluding the subject) to obtain the background information of the context image. Here, the subject refers to the subject contained in the context image, which may be the same as or different from the subject contained in the image to be processed; this is not limited here. Subject information refers to the subject category and location information corresponding to the subject identified in the context image. Since the subject may perform some actions, the subject category may change, and the location information may be the same as or different from the location information of the same subject in the image to be processed; the specific determination depends on the actual situation.
[0071] It should be noted that if the time difference between the context image and the image to be processed meets the time threshold (less than the time threshold), the subject information and background information of the context image can be obtained. The steps for obtaining the subject information of the preceding and following images in the context image are the same as in step S210, and the steps for obtaining the background information of the preceding and following images are the same as in step S220, so they will not be described again. If there is no context image or the time difference between the context image and the image to be processed does not meet the time threshold (greater than the time threshold), the context image analysis ends.
[0072] For example, Table 2 schematically illustrates the subject information of the above images, which may include the subject name and the location information of each subject.
[0073] Table 2
[0074]
[0075] Table 3 schematically illustrates the subject information for the images below. The subject information may include the subject name and the location information of each subject. The subject names represented by the subject categories in the preceding and following images may be the same or different, and the location information of the same subject may be the same or different.
[0076] Table 3
[0077]
[0078] Next, after obtaining the background information of the context image, it can be compared with the background information of the image to be processed to determine whether they are the same. If they are the same, the context image is considered to meet the analysis conditions, and the subject in the image to be processed can be classified and detected based on the context image to obtain the subject's detailed information, thus achieving auxiliary analysis of the detailed information; if they are different, the context image is considered not to meet the analysis conditions, and the context image analysis ends.
[0079] Figure 5 This schematically illustrates a flowchart for classifying and detecting subjects in an image to be processed, incorporating contextual images. (Refer to...) Figure 5 As shown, the main steps include:
[0080] In step S510, the subject of the image to be processed is classified and detected to obtain detailed information about the subject.
[0081] In this embodiment of the disclosure, a classification algorithm can be used to classify and identify all subjects segmented from the image to be processed, and to determine the detailed information of each subject. The detailed information may include positional relationship information and attribute information.
[0082] Table 4
[0083]
[0084] The positional relationship information can be determined based on the positional information of different subjects. Specifically, it can be determined by the positional vector between the top-left corner coordinates, bottom-right corner coordinates, or center coordinates of all subjects. Positional relationship information can represent the position and direction between different subjects, i.e., orientation. Positions can include, but are not limited to, up, down, left, right, intersecting, etc., and directions can be, for example, +30°, +210°. In addition, positional relationship information can also include distances (e.g., 20 pixels, 50 pixels), as shown in Table 4.
[0085] Attribute information can be used to describe the state attributes of each subject, specifically including but not limited to actions, colors, expressions, and other attribute information. The number of actions can be at least one, and actions can be individual actions for each subject or related actions for each subject in relation to other subjects. For example, when the subject is a person, actions can include individual actions such as walking; it can also include related actions between the subject and other subjects, such as dogs, with the related action being walking the dog. For human subjects, facial expressions can be identified. If the subject belongs to other categories, such as animals, buildings, vehicles, etc., facial expressions do not need to be identified. Therefore, the type of attribute information for a subject is not fixed but can be adjusted in real time according to the subject category. Other attribute information can include, but is not limited to, specifications, volume, size, and all other features related to the subject, determined according to actual needs, and is not specifically limited here.
[0086] For example, a machine learning model can be used to classify and detect the subject in the image to be processed, and determine the subject's actions and attribute information such as expression and color, as shown in Table 5.
[0087] Table 5
[0088]
[0089] In step S520, reference detail information of the subject in the context image is obtained.
[0090] In this step, if it is necessary to combine the context image to assist in the analysis of the image to be processed, it is necessary to obtain reference detail information for each subject in the context image. The reference detail information may include positional relationship information and attribute information. The specific method for obtaining this information is the same as that in step S510, and will not be repeated here. See Table 6 for the positional relationship information and Table 7 for the attribute information.
[0091] Table 6
[0092]
[0093] It should be noted that the reference detail information obtained from the context image can be the same as or different from the detail information of the subject in the image to be processed. For example, as shown in Table 5, the action of subject 1 in the image to be processed is running, and the action of subject 1 in the context image in Table 7 is also running.
[0094] Table 7
[0095]
[0096] In step S530, the detail information of each subject is updated based on the reference detail information of the context image and the detail information of the image to be processed.
[0097] In this embodiment of the disclosure, to improve the accuracy and rationality of the detailed information, the detailed information of the subject can be updated using reference detailed information from the context image and the detailed information of the image to be processed, thereby comprehensively determining the final detailed information of each subject in the image to be processed. The detailed information may include positional relationship information and attribute information; the processing method for attribute information will be described in detail here first.
[0098] For example, reference detail information and detail information can be combined using intersection and union calculations. Specifically, the intersection result of reference detail information and the detail information can be obtained. For instance, the intersection of reference detail information and detail information in the preceding image is calculated, and the intersection of reference detail information and detail information in the following image is also calculated to obtain the intersection result.
[0099] Next, the reference detail information and the union of the detail information can be obtained by merging all intersection results to update the detail information of the subject. For example, the detail information can be obtained by taking the union of all intersection results. For instance, the attribute information in the detail information can be determined by (the attribute of the image to be processed ∩ the attribute of the preceding image) ∪ (the attribute of the image to be processed ∩ the attribute of the following image) to update the attribute information in the detail information.
[0100] Based on the above method, if the intersection result determines that the reference detail information is partially similar to the detail information of the same subject in the image to be processed, then the target information between the reference detail information and the detail information is obtained, and the detail information is supplemented according to the target information to obtain a union result, thereby updating the obtained detail information of the subject. Partial similarity means that the similarity is less than a preset value, such as 90%. The target information can be all different information between the detail information and the reference detail information, or it can be different information for the same attribute. That is, the detail information of the image to be processed can be supplemented according to the target information, thereby updating the detail information of the subject. If the intersection result determines that the reference detail information is similar to the detail information, this similarity can be understood as a similarity greater than a preset value. In this case, the reference detail information or the detail information can be directly determined as the detail information of the subject.
[0101] Regarding the positional relationship information in the detailed information, it can be combined with the positional relationship information of the subject in the context image to verify and confirm the positional relationship information of the subject in the image to be processed and the actions contained in the subject's attribute information. Specifically, by comparing the positional relationship information, distance changes, and directions of distance changes between the image to be processed and the context image, the positional change relationship of the subject in the image to be processed can be further determined, and the actions of the subject can be confirmed. For example, as can be seen from Tables 4 and 6, the distance between subject 1 and subject 2 has changed, indicating that both subject 1 and subject 2 have moved.
[0102] For example, if the attribute information of a subject in the context image is completely consistent (similar) to the attribute information of the same subject in the image to be processed, the attribute information of that subject in the image to be processed can be further confirmed. For instance, if subject 1 in both the context image and the image to be processed is smiling, it can be confirmed that subject 1 in the image to be processed has a happy expression. If the attribute information of subject 1 in the context image is not completely consistent (partially similar) to the attribute information of subject 1 in the image to be processed, the target information represented by the different information of subject 1 is extracted as supplementary information to the attribute information of subject 1 in the image to be processed. If the attribute information of subject 1 in the context image is completely different from the attribute information of subject 1 in the image to be processed, then the context image is considered to be unable to provide assistance and has no reference value, and the attribute information of subject 1 in the image to be processed can be directly used as detailed information.
[0103] It should be noted that each subject in the image to be processed can be processed in the manner of steps S510 to S530 to obtain the detailed information of each subject. By combining the reference detailed information of the subject in the context image, the updated detailed information of each subject can be more accurate, comprehensive and complete.
[0104] It should be added that if the context image does not meet the analysis conditions, the detailed information of each subject is determined directly based on the image to be processed itself. The specific steps are the same as in step S510, and will not be repeated here.
[0105] Next, continue to refer to Figure 2 As shown, in step S240, a target subject is determined based on each of the subjects, and semantic information of the image to be processed is generated according to the subject information, the detail information, and the background information of the target subject.
[0106] In this embodiment of the disclosure, since there may be at least one subject in the image to be processed, the description of the subject in the image may be unfocused, leading to a confusing description. Therefore, it is necessary to select a subject as the target subject to achieve an accurate description. In this embodiment of the disclosure, since the subjects are separated, during playback, one subject can be freely selected as the target subject for targeted processing, so as to describe the content of the image to be processed in a targeted manner based on the target subject.
[0107] The target subject can be any one of at least one subjects, specifically determined by the user's selection operation or automatically based on subject priority information. For example, if a selection operation is detected, the target subject can be determined in response to the selection operation. The selection operation can be touch-based, voice-based, etc., without limitation, as long as selection is possible. In addition, the target subject can be switched through switching methods. Switching methods can be, for example, swiping or clicking, etc., as long as the selected subject can be changed.
[0108] refer to Figure 6 As shown, subject recognition is performed on the image 601 to be processed to obtain subjects 602 and 603. The object bounding box 604 corresponding to subject 602 and the object bounding box 605 corresponding to subject 603 are also determined. If a selection operation is detected on object bounding box 604, subject 602 can be identified as the target subject. Furthermore, the target subject can be switched by changing the position of the selection operation. For example, if the position of the selection operation changes from acting on object bounding box 604 to acting on object bounding box 605, the target subject is switched from subject 602 to subject 603.
[0109] In addition, if no selection operation is detected, the target subject can be automatically determined based on the subject's priority information. Specifically, the priority information of all subjects identified in the image to be processed can be determined, and the target subject can be determined based on this priority information. For example, the subject with the highest priority information can be determined as the target subject, or the subject with the lowest priority information can be determined as the target subject, and so on. It should be noted that the priority information of each subject can be configured in advance, or determined based on the proportion of the subject in the entire image to be processed; this is not limited here. For example, if the priority information of the subjects is arranged from highest to lowest as follows: people, animals, buildings, objects, plants, etc., then people are selected as the target subject. As another example, if people occupy the largest proportion in the entire image to be processed, then people are selected as the target subject.
[0110] In this embodiment, by splitting multiple subjects and determining a target subject to generate semantic information for that target subject to determine the descriptive information of the image to be processed, the unfocused and confusing descriptions caused by the absence of a target subject in related technologies are avoided. This improves targeting, accuracy, interpretability, and readability of the descriptive information. Furthermore, the target subject can be switched and changed according to the switching method, and the descriptive information corresponding to the image to be processed can be generated by freely selecting the target subject, achieving convenience, increasing flexibility, and expanding the scope of descriptive information.
[0111] After the target subject is obtained, since the target subject may not exist alone in the image to be processed, but is associated with other reference subjects, reference subjects associated with the target subject can be identified from all subjects in the image to be processed. The number of reference subjects can be zero, one, two, etc., depending on the configuration requirements. For example, when a person is the target subject, the reference subjects could be a horse, a car, a dog, etc.
[0112] In some embodiments, reference subjects associated with the target subject can be determined based on the positional relationship information between the target subject and different subjects. Specifically, since the positional relationship information can include the distance between different subjects, reference subjects can be determined according to the distances between multiple subjects and the target subject. For example, subjects whose distances meet certain distance criteria can be used as reference subjects. Distance criteria could be, for example, minimum distance or distance less than a distance threshold, etc. It should be noted that the distance between the target subject and the reference subject is negatively correlated with the degree of association between them. That is, the smaller the distance, the greater the degree of association between the target subject and the reference subject. For example, referring to Table 4, the distance between subject 1 and subject 2 is less than the distance between subject 1 and subject n; therefore, the degree of association between subject 1 and subject 2 is greater than the degree of association between subject 1 and subject n. Figure 6 As shown, subject 602 in the image to be processed 601 can be taken as the target subject, and subject 603 can be taken as the reference subject. That is, the target subject is a person, and the reference subject is a horse.
[0113] After obtaining the target subject and the reference subject, generating descriptive information for the image to be processed based on the subject information, detailed information, and background information of the target subject can be understood as combining the subject information, attribute information, background information, and detailed information of the reference subject to obtain the descriptive information. Specifically, at least some attribute information from the subject information of the target subject, including the subject category, attribute information, and background information, and from the detailed information of the reference subject, can be combined to obtain descriptive information representing the content of the image to be processed. At least some attribute information of the reference subject may include one or more attributes such as color and expression. Alternatively, it may not include the attribute information of the reference subject, depending on the configured description type and description requirements. Description requirements can indicate the completeness of the description, such as whether it includes partial or complete attribute information of the reference subject.
[0114] It should be noted that when generating descriptive information, the descriptive information can be determined based on the descriptive type. Descriptive types can include, but are not limited to, action-based, positional relationship-based, and hybrid types, depending on the actual configuration. Action-based descriptive information refers to generating descriptive information based on actions to describe the content of the image to be processed. Positional relationship-based descriptive information refers to generating descriptive information based on position. Hybrid types refer to generating descriptive information based on all states (e.g., actions, positions, etc.). In this embodiment, the action-based descriptive type is used as an example for illustration.
[0115] Figure 7 The flowchart for obtaining descriptive information is illustrated in the diagram. (Refer to...) Figure 7 As shown, the main steps include:
[0116] In step S710, the target action of the target subject is determined according to the action weight of at least one action.
[0117] In this step, the target subject's attribute information may include at least one action and other attribute information. The number of actions can be determined according to actual needs, for example, based on the target subject's own actions and related actions with the reference subject. Other attribute information may include color, expression, etc.
[0118] Because misidentification of actions can occur during action recognition of a subject, such as misidentifying a waving hand as a handshake, this can lead to incorrect descriptive information. Therefore, the descriptive information of the image to be processed can be determined based on the action weight of each action identified for the target subject. Multiple actions can be identified for each target subject, each with a different action weight. The higher the action weight, the stronger the correlation between the action and the target subject. For example, the target action of the target subject can be determined based on the action weight. Specifically, the action with the highest action weight can be identified as the target action of the target subject. In some embodiments, all actions with weights greater than a preset value can also be identified as the target actions of the target subject. Determining the target action by action weight avoids the problem of potentially using incorrectly identified actions to construct descriptive information, thus improving the accuracy of identifying the target action of the subject and consequently improving the accuracy of the constructed descriptive information, making the descriptive information more consistent with reality and thus enhancing the rationality of the descriptive information.
[0119] For example, it can be done by Figure 7 In step S710, the target action is selected from action 1, action 2, and action 3. (See reference...) Figure 8 As shown, the target subject is the person represented by subject 801, the reference subject is the dog represented by subject 802, and the car represented by subject 803. Among the actions identified for the person, the weight of the sitting action is greater than the weight of the standing action. Based on the action weight, it can be determined that the action of the target subject may include the person sitting in the car, rather than standing in the car.
[0120] In step S720, according to language rules, the subject information of the target subject, the target action, other attribute information, the background information, and at least some attribute information of the reference subject are combined into the description information.
[0121] In this step, the language rules can be grammatical rules. The grammatical rules can be determined according to actual needs. For example, grammatical rules can be subject-verb-object, or subject-verb-object-predicate-complement, etc. Here, we will take the subject-verb-object rule as an example for explanation.
[0122] In some embodiments, after obtaining all descriptive information of the target subject, step S720 can be used to combine the subject information, target action, other attribute information, background information, and at least some attribute information of the reference subject according to the subject-verb-object rule to form one or more sentences to represent the descriptive information of the image to be processed, and the descriptive information is described from the dimension of the target subject. Based on this, the subject can be the target subject, the predicate can be the action of the target subject, etc., without specific limitations here.
[0123] For example, refer to Figure 6As shown, the extracted descriptive information may include, for example, the following:
[0124] Background information: Grassland; Subject information: Person, horse. Positional relationship between subjects: The person and the horse intersect; the person is on the horse, indicating riding. Subject attribute information: Person (white clothes, hat, smiling, cheerful), Horse (black, walking, standing). If the person is the target subject, the description information could be: A person is happily riding a horse on the grass. If the description information needs to be generated based on all attribute information of all subjects, it could also be represented as: A person wearing a hat and white clothes is happily riding a black horse on the grass. Alternatively, the target subject can be changed to a horse, and the description information could be: A black horse is happily moving on the grass, with a person wearing a hat and white clothes riding on its back. Based on this, description information generated based on all information for a specific target subject can comprehensively and completely describe the content of the entire image to be processed, improving accuracy.
[0125] In this embodiment, by configuring action weights to generate descriptive information that matches the content of the image to be processed, the rationality and accuracy of the descriptive information can be improved, as well as the matching between the descriptive information and the image to be processed. Furthermore, because it combines the attribute information of the target subject, background information, and at least some attribute information of the reference subject, compared to related technologies that only rely on single-point information, the completeness and comprehensiveness of the descriptive information are improved, generating descriptive information from multiple dimensions to enhance its accuracy.
[0126] This disclosure also provides an image description method, referring to... Figure 9 As shown, the main steps include:
[0127] In step S910, an image to be processed is acquired, and descriptive information of the image to be processed is generated; the descriptive information is generated based on the subject information, detail information, and background information of the target subject in the image to be processed.
[0128] In step S920, the description information is played to read the image to be processed.
[0129] In this embodiment, the descriptive information is used to represent the semantic information of the content of the image to be processed. It can be generated based on the subject information and detail information of the target subject contained in the image to be processed, as well as the background information of the image. The subject information may include the subject category and the subject's location information; the detail information may include positional relationship information and attribute information, and the attribute information may include actions and other attribute information; the background information may be the background category. The generation methods of the subject information, detail information, background information, and descriptive information are the same as in steps S210 to S240, and will not be repeated here.
[0130] In this embodiment of the disclosure, after obtaining the description information, if the user may be unable to view the image or the image cannot be displayed due to network problems, the description information can be played. The playback method can include text or voice, depending on the display state of the image to be processed. For example, if the display state of the image to be processed is normal, the playback method can be voice; if the display state of the image to be processed is abnormal, the playback method can be either text or voice.
[0131] Next, we will take voice playback as an example. For instance, to facilitate voice playback, the descriptive information corresponding to the generated image to be processed can be converted into voice information and played back. This allows users to quickly learn the content of the image to be processed in scenarios where it is inconvenient to view the image, thus enabling barrier-free image reading through voice.
[0132] When playing descriptive information via voice, the playback can be based on the current scene category and device status. The current scene refers to the current scene in which the user viewing the image to be processed is located. The current scene can be any scene, and its category can be either a playable scene or not. Playable scenes include, but are not limited to, private scenes and scenes where no other content is playing; scenes not playing include, but are not limited to, public scenes and scenes where other content is playing (e.g., playing music). Device status can be the receiving device status, where the receiving device is a playback device connected to the terminal displaying the image to be processed, such as headphones. The receiving device status can also be the connection status of the receiving device, i.e., whether the receiving device is connected to the terminal displaying the image to be processed.
[0133] In some embodiments, if the scene category does not belong to a playable scene and the receiving device is in an unconnected state, the description information is played in response to a playback command and according to playback parameters; if the scene category belongs to a playable scene or the receiving device is in a connected state, the description information is played according to the corresponding playback parameters.
[0134] Specifically, if the current scene category does not belong to a playable scene, the receiving device status is further determined. If the receiving device status is not connected, the description information is controlled to play in response to the playback command. The playback command can be a voice command or a playback command executed in other forms; here, a voice command is used as an example. For example, the playback command can be the voice command "Play". Playback parameters can be used to indicate specific playback details, including but not limited to playback volume, number of plays, etc. Different scene types have different corresponding playback parameters.
[0135] Therefore, if the current scene category does not belong to a playable scene and the receiving device is in an unconnected state, if a playback command is received, the description information can be played according to the playback parameters corresponding to the current scene type to achieve barrier-free reading of the image to be processed. If no playback command is received, playback will not occur.
[0136] If the current scene category does not belong to a playable scene and the receiving device is in a connected state, the description information can be played directly according to the playback parameters corresponding to the current scene type to achieve barrier-free reading of the image to be processed without needing to determine whether a playback command has been received.
[0137] If the current scene category is a playable scene, there is no need to continue judging the receiving device status. The description information can be played directly according to the playback parameters corresponding to the current scene type to achieve barrier-free reading of the image to be processed.
[0138] In this embodiment, by playing descriptive information generated based on the subject's information, details, and background information, the convenience for users to view the image to be processed is improved, enhancing the user experience and expanding the application scope. Furthermore, it avoids the problem of incomplete descriptive information, providing users with comprehensive, complete, and accurate descriptive information about the image to be processed. By combining the scene category of the current scene and the status of the receiving device to determine the playback method of the descriptive information, interference or inappropriate situations caused by direct playback can be avoided, improving privacy and rationality.
[0139] It should be noted that the solution provided in this embodiment can also convert the content contained in each frame of the video into descriptive information for playback. The specific method is the same as the steps for playing the descriptive information of the image, and will not be repeated here.
[0140] Figure 10 The diagram illustrates the specific flowchart for reading the image to be processed. (See reference...) Figure 10 As shown, the main steps include:
[0141] In step S1001, the image to be processed is acquired.
[0142] In step S1002, it is determined whether the time difference between the context image and the image to be processed is less than the time threshold; if yes, proceed to step S1004; if no, end the context image analysis.
[0143] In step S1003, subject recognition is performed on the image to be processed to obtain subject information.
[0144] In step S1004, subject recognition is performed on the context image to obtain subject information.
[0145] In step S1005, background recognition is performed on the image to be processed to obtain background information. For example, subject removal and background padding can be performed to obtain background information.
[0146] In step S1006, background recognition is performed on the context image to obtain background information.
[0147] In step S1007, it is determined whether the background information of the image to be processed is consistent with the background information of the context image. If yes, proceed to step S1009; otherwise, end the context image analysis.
[0148] In step S1008, subject detail analysis is performed on the subject in the image to be processed to obtain detail information. The detail information may include positional relationship information and attribute information, and the attribute information may include action and other attribute information.
[0149] In step S1009, subject detail analysis is performed on the subject in the context image to obtain reference detail information.
[0150] In step S1010, the detail information of the image to be processed is updated based on the reference detail information of the context image as supplementary information to obtain the detail information of the subject.
[0151] In step S1011, the target subject is selected.
[0152] In step S1012, descriptive information describing the target subject in the image to be processed is organized.
[0153] In step S1013, the description information is played aloud.
[0154] The technical solution in this disclosure generates descriptive information from multiple dimensions—combining the subject information, detail information, and background information of the image to be processed—to describe the content of the image. This descriptive information is then played to enable image reading. Compared to existing technologies, this avoids the limitation of generating descriptive information based only on partial information, increasing comprehensiveness, improving the accuracy of the descriptive information, and enhancing the matching and rationality between the descriptive information and the content of the image to be processed, thus improving realism. Because a target subject can be selected to generate the descriptive information, the problem of inaccurate descriptions in related technologies is avoided, improving the accuracy and relevance of the descriptive information. Furthermore, playing the descriptive information describing the image to be processed can be applied to any scenario where viewing the image is inconvenient or impossible, expanding the application scope and improving convenience.
[0155] This disclosure provides an image description device, with reference to... Figure 11 As shown, the image description device 1100 may include:
[0156] The subject information acquisition module 1101 is used to acquire the image to be processed, perform target detection on the image to be processed, and determine the subject information of at least one subject contained in the image to be processed;
[0157] Background information acquisition module 1102 is used to acquire background information based on the remaining image in the image to be processed, excluding each of the main subjects;
[0158] The detail information determination module 1103 is used to classify and detect the main body in the image to be processed, and determine the detail information of each main body.
[0159] The description information generation module 1104 is used to determine the target subject based on each of the subjects, and generate description information of the image to be processed according to the subject information, the detail information, and the background information of the target subject.
[0160] In one exemplary embodiment of this disclosure, the subject information acquisition module includes: a subject recognition module, configured to determine the subject category of all subjects in the image to be processed through a target detection model, and acquire the location information of each subject.
[0161] In one exemplary embodiment of this disclosure, the background information acquisition module includes: a background completion module, used to extract each of the subjects to obtain the remaining image, and to perform background completion on the remaining image to obtain the completed image; and a category recognition module, used to recognize the completed image through a target detection model to obtain the background category as background information.
[0162] In one exemplary embodiment of this disclosure, the detail information determination module includes: a comprehensive determination module, configured to acquire a context image of the image to be processed, classify and detect subjects in the image to be processed in conjunction with the context image, and determine the detail information of each subject.
[0163] In one exemplary embodiment of this disclosure, the comprehensive determination module includes: a detail information determination module, used to classify and detect the subject of the image to be processed and obtain the detail information of the subject; a reference detail information determination module, used to obtain reference detail information of the subject in the context image; and an information fusion module, used to update the detail information of each subject according to the reference detail information of the context image and the detail information of the image to be processed.
[0164] In one exemplary embodiment of this disclosure, the information fusion module includes: an intersection result determination module, configured to obtain the intersection result of the reference detail information and the detail information; and a union result determination module, configured to merge the intersection result to obtain the union result of the reference detail information and the detail information, so as to update the detail information.
[0165] In one exemplary embodiment of this disclosure, the union result determination module includes: a first determination module, configured to, if determined by the intersection result that the reference detail information is partially similar to the detail information, obtain target information between the reference detail information and the detail information, and supplement the detail information according to the target information to obtain a union result, thereby obtaining the detail information of the subject; and a second determination module, configured to, if determined by the intersection result that the reference detail information is similar to the detail information, determine the reference detail information or the detail information as the detail information of the subject.
[0166] In an exemplary embodiment of this disclosure, the detailed information includes positional relationship information and attribute information; the description information generation module includes: a reference subject determination module, used to obtain a reference subject associated with the target subject based on the positional relationship information; and an information combination module, used to combine the subject information, attribute information, background information of the target subject, and detailed information of the reference subject to obtain the description information.
[0167] In one exemplary embodiment of this disclosure, the reference subject determination module includes: a subject selection module, configured to determine the reference subject based on the distances between multiple subjects and the target subject in the positional relationship information; the distances are negatively correlated with the degree of association of the reference subject.
[0168] In one exemplary embodiment of this disclosure, the information combination module includes: an information generation module, configured to obtain a description type and combine at least a portion of the attribute information from the subject information, attribute information, background information, and detailed information of the reference subject of the target subject into description information that satisfies the description type.
[0169] In one exemplary embodiment of this disclosure, the attribute information includes at least one action and other attribute information; the information generation module includes: a target action determination module, configured to determine the target action of the target subject based on the action weight of at least one action; and a generation control module, configured to combine the subject information of the target subject, the target action, other attribute information, background information, and at least some attribute information of the reference subject into the description information according to language rules.
[0170] In an exemplary embodiment of this disclosure, before obtaining the reference detail information of the context image, the apparatus further includes: a subject information determination module, configured to perform subject recognition on the context image if a context image exists and the time difference between the image to be processed and the context image is less than a time threshold, and obtain subject information of the subject in the context image; and a background information determination module, configured to perform target detection on the remaining images in the context image other than the subject, and obtain background information of the context image.
[0171] In one exemplary embodiment of this disclosure, the target subject determination module includes: a selection module, configured to determine the subject corresponding to the selection operation as the target subject in response to a selection operation; or a default selection module, configured to obtain priority information of each subject and determine the target subject based on the priority information.
[0172] This disclosure provides another image description device, with reference to... Figure 12 As shown, the image description device 1200 may include:
[0173] The description information generation module 1201 is used to acquire the image to be processed and generate description information of the image to be processed; the description information is generated based on the subject information, detail information and background information of the target subject in the image to be processed.
[0174] The information playback module 1202 is used to play the description information in order to read the image to be processed.
[0175] In one exemplary embodiment of this disclosure, the information playback module includes a playback control module, used to play the descriptive information via text or voice.
[0176] It should be noted that the specific details of each part of the above-mentioned image description device have been described in detail in the implementation of the image description method. For any undisclosed details, please refer to the implementation of the method section, and therefore will not be repeated here.
[0177] Exemplary embodiments of this disclosure also provide an electronic device. This electronic device may be the terminal 101 described above. Generally, the electronic device may include a processor and a memory, the memory being used to store executable instructions of the processor, the processor being configured to perform the image description method described above by executing the executable instructions.
[0178] The following is based on Figure 13 Taking the mobile terminal 1300 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 13 The structure can also be applied to fixed types of equipment.
[0179] like Figure 13 As shown, the mobile terminal 1300 may specifically include: a processor 1301, a memory 1302, a bus 1303, a mobile communication module 1304, an antenna 1, a wireless communication module 1305, an antenna 2, a display screen 1306, a camera module 1307, an audio module 1308, a power module 1309, and a sensor module 1310.
[0180] Processor 1301 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). In this exemplary embodiment, the method can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU. For example, the NPU can load neural network parameters and execute neural network-related algorithm instructions.
[0181] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 1300 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0182] The processor 1301 can be connected to the memory 1302 or other components via the bus 1303.
[0183] The memory 1302 can be used to store computer executable program code, which includes instructions. The processor 1301 executes various functional applications and data processing of the mobile terminal 1300 by running the instructions stored in the memory 1302. The memory 1302 can also store application data, such as images, videos, and other files.
[0184] The communication function of mobile terminal 1300 can be implemented through mobile communication module 1304, antenna 1, wireless communication module 1305, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1304 can provide 3G, 4G, and 5G mobile communication solutions for mobile terminal 1300. Wireless communication module 1305 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1300.
[0185] The display screen 1306 is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module 1307 is used to implement shooting functions, such as capturing images and videos, and may include a color temperature sensor array. The audio module 1308 is used to implement audio functions, such as playing audio and capturing voice. The power module 1309 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1310 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 1310 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1300 and output inertial sensing data.
[0186] It should be noted that the present disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist alone and not assembled into the electronic device.
[0187] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0188] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0189] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments.
[0190] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0191] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0192] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0193] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims. It should be understood that this disclosure is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image description method, characterized in that, include: Acquire an image to be processed, and perform target detection on the image to be processed to determine the subject information of at least one subject contained in the image to be processed; Background information is obtained from the remaining image in the image to be processed, excluding the main subjects. The main subjects in the image to be processed are classified and detected to determine the detailed information of each subject; Based on each of the aforementioned subjects, a target subject is determined, and descriptive information of the image to be processed is generated according to the subject information, detailed information, and background information of the target subject. The step of classifying and detecting the subjects in the image to be processed, and determining the detailed information of each subject, includes: The system classifies and identifies the main subject in the image to be processed, determines the detailed information of the subject, and obtains reference detailed information of each subject in the context image. Based on the intersection result of the reference detail information and the detail information, it is determined that the reference detail information is partially similar to the detail information. The target information between the reference detail information and the detail information is obtained, and the detail information is supplemented according to the target information to obtain the union result, so as to obtain the detail information of the subject. If the reference detail information is determined to be similar to the detail information based on the intersection result, the reference detail information or the detail information is determined to be the detail information of the subject.
2. The image description method according to claim 1, characterized in that, The step of performing target detection on the image to be processed to determine the subject information of at least one subject contained in the image to be processed includes: The object detection model is used to determine the object category of all objects in the image to be processed, and the location information of each object is obtained.
3. The image description method according to claim 1, characterized in that, The step of obtaining background information from the remaining image in the image to be processed, excluding the subjects, includes: Each of the aforementioned subjects is extracted to obtain the remaining image, and the background of the remaining image is filled in to obtain the filled image; The imputed image is identified using an object detection model to obtain the background category as background information.
4. The image description method according to claim 1, characterized in that, The step of classifying and detecting the subjects in the image to be processed, and determining the detailed information of each subject, includes: Obtain the context image of the image to be processed, and combine the context image to classify and detect the subjects in the image to be processed, thereby determining the detailed information of each subject.
5. The image description method according to claim 4, characterized in that, The step of classifying and detecting subjects in the image to be processed by combining the context image, and determining the detailed information of each subject, includes: The detailed information of each subject is updated based on the reference detail information of the context image and the detail information of the image to be processed.
6. The image description method according to claim 5, characterized in that, The step of updating the detail information of each subject based on the reference detail information of the subject in the context image and the detail information of the image to be processed includes: Obtain the intersection result of the reference detail information and the detail information; The intersection results are merged to obtain the union result of the reference detail information and the detail information, so as to update the detail information.
7. The image description method according to claim 1, characterized in that, The detailed information includes location relationship information and attribute information; The step of generating descriptive information for the image to be processed based on the subject information, the detail information, and the background information of the target subject includes: Based on the location relationship information, obtain the reference subject associated with the target subject; The description information is obtained by combining the subject information, attribute information, background information of the target subject, and detailed information of the reference subject.
8. The image description method according to claim 7, characterized in that, The step of obtaining the reference subject associated with the target subject based on the location relationship information includes: The reference subject is determined based on the distances between multiple subjects and the target subject in the location relationship information; the distances are negatively correlated with the degree of association between the reference subject and the target subject.
9. The image description method according to claim 7, characterized in that, The step of combining the subject information, attribute information, background information, and detailed information of the reference subject of the target subject to obtain the descriptive information includes: Obtain the description type, and combine at least some of the attribute information from the subject information, attribute information, background information, and detailed information of the reference subject of the target subject into description information that satisfies the description type.
10. The image description method according to claim 9, characterized in that, The attribute information includes at least one action and other attribute information; The step of combining at least a portion of the attribute information from the subject information, attribute information, background information, and detailed information of the reference subject into descriptive information that satisfies the description type includes: The target action of the target subject is determined based on the action weight of at least one action; According to language rules, the subject information, target action, other attribute information, background information, and at least some attribute information of the reference subject are combined to form the description information.
11. The image description method according to claim 5, characterized in that, Before obtaining the reference detail information of the context image, the method further includes: If a context image exists and the time difference between the image to be processed and the context image is less than a time threshold, subject recognition is performed on the context image to obtain subject information in the context image. Target detection is performed on the remaining images in the context image, excluding the main subject, to obtain the background information of the context image.
12. The image description method according to claim 1, characterized in that, The determination of the target subject based on each of the aforementioned subjects includes: In response to a selection operation, the subject corresponding to the selection operation is determined as the target subject; or Obtain the priority information of each subject, and determine the target subject based on the priority information.
13. An image description method, characterized in that, include: Acquire the image to be processed and generate description information for the image to be processed; The description information is generated based on the subject information, detail information, and background information of the target subject in the image to be processed; The description information is played to read the image to be processed; The methods for determining the detailed information include: The system classifies and identifies the main subject in the image to be processed, determines the detailed information of the subject, and obtains reference detailed information of each subject in the context image. Based on the intersection result of the reference detail information and the detail information, it is determined that the reference detail information is partially similar to the detail information. The target information between the reference detail information and the detail information is obtained, and the detail information is supplemented according to the target information to obtain the union result, so as to obtain the detail information of the subject. If the reference detail information is determined to be similar to the detail information based on the intersection result, the reference detail information or the detail information is determined to be the detail information of the subject.
14. The image description method according to claim 13, characterized in that, Playing the description information includes: The description information can be played via text or voice.
15. An image description device, characterized in that, include: The subject information acquisition module is used to acquire the image to be processed, perform target detection on the image to be processed, and determine the subject information of at least one subject contained in the image to be processed. The background information acquisition module is used to acquire background information based on the remaining image in the image to be processed, excluding each of the main subjects. The detail information determination module is used to classify and detect the main subjects in the image to be processed, and determine the detail information of each subject. The description information generation module is used to determine the target subject based on each of the subjects, and generate description information of the image to be processed according to the subject information, the detail information and the background information of the target subject; The step of classifying and detecting the subjects in the image to be processed, and determining the detailed information of each subject, includes: The system classifies and identifies the main subject in the image to be processed, determines the detailed information of the subject, and obtains reference detailed information of each subject in the context image. Based on the intersection result of the reference detail information and the detail information, it is determined that the reference detail information is partially similar to the detail information. The target information between the reference detail information and the detail information is obtained, and the detail information is supplemented according to the target information to obtain the union result, so as to obtain the detail information of the subject. If the reference detail information is determined to be similar to the detail information based on the intersection result, the reference detail information or the detail information is determined to be the detail information of the subject.
16. An image description device, characterized in that, include: The description information generation module is used to acquire the image to be processed and generate description information of the image to be processed; The description information is generated based on the subject information, detail information, and background information of the target subject in the image to be processed; An information playback module is used to play the description information in order to read the image to be processed; The step of classifying and detecting the subjects in the image to be processed, and determining the detailed information of each subject, includes: The system classifies and identifies the main subject in the image to be processed, determines the detailed information of the subject, and obtains reference detailed information of each subject in the context image. Based on the intersection result of the reference detail information and the detail information, it is determined that the reference detail information is partially similar to the detail information. The target information between the reference detail information and the detail information is obtained, and the detail information is supplemented according to the target information to obtain the union result, so as to obtain the detail information of the subject. If the reference detail information is determined to be similar to the detail information based on the intersection result, the reference detail information or the detail information is determined to be the detail information of the subject.
17. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the image description method according to any one of claims 1-14 by executing the executable instructions.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image description method according to any one of claims 1-14.
Citation Information
Patent Citations
Image processing method and device, and computer storage medium
CN110933299A
Method and apparatus for commenting video
US20210357653A1