Image processing method, device, apparatus, and computer-readable storage medium

By acquiring image information of the target object and using machine learning models to identify the environment type, the technical gap in urban color analysis and environmental diagnosis is solved, and high-precision environmental recognition is achieved.

CN116091921BActive Publication Date: 2025-10-03HANGZHOU LIFEI SOFTWARE TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211709203.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2025-10-03
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

Existing technologies fail to effectively implement urban color analysis and environmental diagnosis, and lack relevant image recognition technology.

Method used

By acquiring images of target objects, extracting their category information, confidence, color category, and color ratio, the trained machine learning model is used to identify the environment type, and environmental changes are judged based on changes in color ratio.

Benefits of technology

Accurately identifying the type of environment an object is in provides a basis for urban color analysis and environmental diagnosis, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091921B_ABST
    Figure CN116091921B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose an image processing method, apparatus, device, and computer-readable storage medium, wherein the method may include the following steps: acquiring at least one image including a target object; wherein each image carries environmental information, and the environmental information includes category information of the target object, confidence level of the target object, color category of the target object, and color ratio; confidence level is used to characterize the credibility of the environmental information of the surrounding environment of the target object; color ratio is used to characterize the area ratio of the color category of the target object in the entire image; based on the category information of the target object, confidence level of the target object, color category of the target object, and color ratio, one or more environment types in which the target object is located in at least one image are determined. By implementing the present application, the environment type in which the object is located can be accurately identified, providing a basis for color analysis and environmental diagnosis of cities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Image recognition refers to the technology of using computers to process, analyze, and understand images to identify targets and objects of various different patterns. It is a practical application of deep learning algorithms and generally includes four steps: image acquisition, image preprocessing, feature extraction, and image recognition. At present, image recognition technology is generally divided into face recognition and product recognition. Face recognition is mainly used in security inspections, identity verification, and mobile payments; product recognition is mainly used in the circulation of goods, especially in unmanned retail areas such as unmanned shelves and smart retail cabinets. The applicant found in the research that no relevant image recognition technology has been proposed to achieve color analysis and environmental diagnosis of cities. Summary of the Invention

[0003] The embodiments of the present application provide an image processing method, apparatus, device, and computer-readable storage medium that can accurately identify the type of environment in which an object is located, providing a basis for color analysis and environmental diagnosis of a city.

[0004] In a first aspect, an embodiment of the present application provides an image processing method, the method comprising:

[0005] Acquire at least one image including a target object; wherein each image carries environmental information, the environmental information including category information of the target object, a confidence level of the target object, a color category of the target object, and a color ratio; the confidence level is used to represent the credibility of the environmental information of the target object's surrounding environment; and the color ratio is used to represent the area ratio of the target object's color category in the entire image;

[0006] One or more environment types in which the target object is located in the at least one image are determined based on the category information of the target object, the confidence level of the target object, the color category and the color proportion of the target object.

[0007] By implementing the embodiments of the present application, the type of environment in which the object contained in the image is located is identified based on the object's category information, the object's confidence, the object's color category, and the color ratio. In this way, the type of environment in which the object is located can be accurately identified, providing a basis for urban color analysis and environmental diagnosis.

[0008] In one possible implementation, determining one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion includes:

[0009] At least one image of the target object is processed by a trained machine learning model to determine one or more types of environments in which the target object in the at least one image is located, wherein the trained machine learning model is a model trained based on at least one sample image, and each sample image is associated with annotation data, and the annotation data includes object category information, object confidence, object color category, and color proportion.

[0010] In one possible implementation, determining one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion includes:

[0011] Processing each of the at least one image separately, and determining a probability distribution of the type of environment in which the target object is located in each image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion;

[0012] The probability distribution corresponding to each of the images is integrated to obtain one or more environment types in which the target object in the at least one image is located.

[0013] In a possible implementation, the imaging quality corresponding to each of the at least one image of the target object is greater than a preset threshold.

[0014] In a possible implementation, the color ratio of the target object is a main color ratio. For example, the color of the target object generally includes multiple color categories. When the environment type of the object is identified based on the main color ratio, the recognition accuracy can be guaranteed.

[0015] In a possible implementation, the similarity between the multiple environment types is greater than a second preset threshold.

[0016] In one possible implementation, the at least one image includes a first image and a second image, the first image carries first time information, and the second image carries second time information, where the second time information is a time point after the first time information; and the method further includes:

[0017] respectively obtaining a color ratio of each color in the first image and the second image;

[0018] When the color ratios of each color corresponding to the first image and the second image are inconsistent, it is determined that the surrounding environment of the target object at the second time information has changed.

[0019] In a second aspect, an embodiment of the present application provides an image processing device, comprising:

[0020] An image acquisition unit, configured to acquire at least one image including a target object; wherein each image carries environmental information, including category information of the target object, a confidence level of the target object, a color category of the target object, and a color ratio; the confidence level is used to represent the credibility of the environmental information of the target object's surrounding environment; and the color ratio is used to represent the area ratio of the target object's color category in the entire image;

[0021] A processing unit is configured to determine one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color ratio.

[0022] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the processor and the memory are connected to each other, wherein the memory is used to store a computer program that supports the electronic device to execute the above-mentioned method, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method of the above-mentioned first aspect.

[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method of the first aspect above.

[0024] In a fifth aspect, an embodiment of the present application further provides a computer program, which includes program instructions, and when the program instructions are executed by a processor, the processor executes the method of the first aspect above.

[0025] The image processing method, apparatus, device, and computer-readable storage medium provided by the embodiments of the present disclosure collect multiple images including a target object, each image carrying environmental information, including the target object's category information, the target object's confidence level, the target object's color category, and the color ratio; the confidence level is used to characterize the credibility of the environmental information surrounding the target object; the color ratio is used to characterize the area ratio of the target object's color category in the entire image; and based on the target object's category information, the target object's confidence level, the target object's color category, and the color ratio, one or more environmental types in at least one image are determined. In this way, the accuracy of identifying the target object's environmental type can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments.

[0027] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0028] Figure 2 This is a flow chart of a machine learning model training method provided in an embodiment of the present application;

[0029] Figure 3 This is a flowchart of an image processing method provided by an embodiment of the present application;

[0030] Figure 4 A flowchart of another image processing method provided in an embodiment of the present application;

[0031] Figure 5 is a structural diagram of an image processing device provided in an embodiment of the present application;

[0032] Figure 6 It is a structural diagram of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0034] The terms "first," "second," "third," and "fourth," etc., in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, rather than to describe a specific order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0035] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0036] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0037] First, refer to Figure 1 To describe an example electronic device 100 for implementing the machine learning model training method and apparatus or image processing method and apparatus according to an embodiment of the present application. Figure 1 As shown, the electronic device 100 includes one or more processors 102, one or more storage devices 104, an input device 106, and an output device 108. The electronic device 100 may also include a data acquisition device 110 and / or an image acquisition device 112, and these components are interconnected via a bus system 114 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1The components and structure of the electronic device 100 shown are merely exemplary and non-limiting. The electronic device may also have other components and structures as needed.

[0038] The processor 102 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.

[0039] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM, hard disk, flash memory), etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may run the program instructions to implement the client functions and / or other desired functions in the embodiments of the present application (implemented by the processor) described below. Various applications and various data may also be stored in the computer-readable storage medium, such as various data used and / or generated by the application.

[0040] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.

[0041] The output device 108 may output various information (eg, images and / or sounds) to the outside (eg, a user), and may include one or more of a display, a speaker, and the like.

[0042] When the electronic device 100 is used to implement the machine learning model training method and apparatus and the image processing method and apparatus according to the embodiments of the present application, the electronic device 100 may include a data acquisition device 110. The data acquisition device 110 may acquire sample images (including video frames) and corresponding annotation data, and store the acquired sample images and annotation data in the storage device 104 for use by other components. For example, the data acquisition device 110 may include one or more of a wired or wireless network interface, a universal serial bus (USB) interface, an optical disk drive, and the like.

[0043] When the electronic device 100 is used to implement the image processing method and apparatus according to the embodiments of the present application, the electronic device 100 may include an image acquisition device 112. The image acquisition device 112 may acquire images (including video frames) and store the acquired images in the storage device 104 for use by other components. The image acquisition device 112 may be a camera. It should be understood that the image acquisition device 112 is merely an example, and the electronic device 100 may not include the image acquisition device 112. In this case, other devices with image acquisition capabilities may be used to acquire images to be processed and transmit the acquired images to the electronic device 100.

[0044] Exemplarily, the example electronic device for implementing the machine learning model training method and apparatus and the image processing method and apparatus according to the embodiments of the present application can be implemented on a device such as a personal computer or a remote server.

[0045] The following combination Figure 2 Specifically introduce the machine learning model training method proposed in this application, such as Figure 2 As shown, the method may include but is not limited to the following steps:

[0046] In step S201 , a sample image and corresponding annotation data are obtained, wherein the annotation data includes object category information, object confidence, object color category, and color proportion.

[0047] The object may be any object, such as a person, a car, a building, etc. In the following description, the present application will be described using a target object as an example, but this is not a limitation of the present application, and the present application may be applicable to matting applications of other objects.

[0048] The sample image can be an original image captured by the electronic device 100 or an image obtained by preprocessing the original image. The ground truth data can be manually annotated. For example, the ground truth data can indicate the category of each pixel in the sample image, and pixels belonging to different objects can be represented by different colors to distinguish them.

[0049] The sample images and annotation data can be sent by a remote device (such as a server storing a training data set) to the electronic device 100 for machine learning model training by the processor 102 of the electronic device 100, or they can be acquired by the data acquisition device 110 included in the electronic device 100 and transmitted to the processor 102 for machine learning model training.

[0050] In step S202, the machine learning model is trained using the loss function, sample images, and corresponding labeled data to obtain a trained machine learning model.

[0051] In the process of training the machine learning model based on the loss function, the back propagation algorithm can be used to adjust the parameters (or weights) used in the machine learning model until the training converges to obtain a trained machine learning model. Exemplarily, the machine learning model training method according to the embodiment of the present application can be implemented in a device, apparatus or system having a memory and a processor. The machine learning model training method according to the embodiment of the present application can be independently deployed on the client or the server. Alternatively, the machine learning model training method according to the embodiment of the present application can also be distributedly deployed on the server (or cloud) and the client. For example, sample images and annotation data can be obtained on the client, and the client transmits the obtained image to the server (or cloud), and the server (or cloud) performs machine learning model training.

[0052] According to embodiments of the present application, the machine learning model can be implemented using a neural network. Of course, the present application is not limited to neural networks; any model based on parameter learning can be applied to the present application. A neural network is a network capable of autonomous learning and has powerful image processing capabilities, making it a very good model choice. In addition, embodiments of the present application can also enhance the sample image in various ways, thereby enhancing the robustness of the machine learning model.

[0053] The machine learning model training method provided in the embodiment of the present application is mainly effective during the training process and does not affect the deployment and use of the machine learning model.

[0054] According to another aspect of the present application, an image processing method is provided. Figure 3 A schematic flowchart of an image processing method according to an embodiment of the present application is shown. The method may include but is not limited to the following steps:

[0055] Step S301: Acquire at least one image including a target object; wherein each image carries environmental information, and the environmental information includes category information of the target object, confidence of the target object, color category of the target object, and color proportion; the confidence is used to characterize the credibility of the environmental information of the surrounding environment of the target object; and the color proportion is used to characterize the area proportion of the color category of the target object in the entire image.

[0056] The image including the target object may be an image captured by an image capture device or an image obtained after preprocessing. The image including the target object may be a static image or a video frame in a video stream.

[0057] An image including a target object can be sent by a client device (such as a mobile terminal including a camera) to the electronic device 100 for image processing by the processor 102 of the electronic device 100, or it can be captured by the image acquisition device 110 included in the electronic device 100 and transmitted to the processor 102 for image processing.

[0058] The color category of the target object may include at least one of white, black, red, yellow, green, blue, orange, brown, and purple.

[0059] It should be noted that in different environmental scenarios, such as different weather conditions (overcast, sunny, snowy, rainy), the surrounding environmental information of the target object is different. In this case, the confidence level may be different in different environmental scenarios. Specifically, the method of obtaining the credibility of environmental information may be different for different environmental information. For example, the probability of an object being determined to be a certain type of object can be used as the confidence level.

[0060] The color ratio of a target object refers to the area ratio of the target object's color category within the entire image. Generally, a target object can include multiple color categories. To improve the recognition accuracy of the target object's environment type, the color ratio can be the dominant color ratio.

[0061] For example, the target object in an image could be green leaves in a park, where the color category is green and the color ratio is 80%. Another example could be a zebra crossing at an intersection, where the color category is white and the color ratio is 80%. Another example could be a billboard on the side of the road, where the color category is blue and the color ratio is 60%.

[0062] For example, in an embodiment of the present application, assume that the image acquisition device experiences jitter while capturing an image of a target object, resulting in lens blur. In this case, misjudgment is likely to occur. When the image quality corresponding to at least one image of the target object is greater than a first preset threshold, the accuracy of identifying the type of environment in which the target object is located can be improved.

[0063] Step S302: Determine one or more environment types in which the target object is located in at least one image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color ratio.

[0064] It is understood that environment types can include weather conditions, site conditions, and lighting conditions. Weather conditions can include sunny, rainy, snowy, and overcast days. Site conditions can include parks, intersections, stores, flat surfaces, pits, and depressions. Lighting conditions can include daytime, nighttime, and strong light. Furthermore, this environment type can also be referred to as a scene type.

[0065] For example, at least one image of the target object may be processed using a trained machine learning model to determine one or more types of environments that the target object is located in. For example, each of 10 images may be input into a trained machine learning model to identify one or more types of environments that the target object is located in.

[0066] For another example, each of at least one image can be processed separately to determine a probability distribution of the target object's environment type in each image based on the target object's category information, the target object's confidence level, the target object's color category, and the color ratio. The probability distributions corresponding to each image are then integrated to obtain one or more environment types for the target object in the at least one image. For example, the prediction result for each of the 10 images is a multidimensional array, which can be a probability distribution. The probability values ​​are then integrated using an integration model. For example, the output of the integration model can be three probability values ​​representing the probabilities of sunny, rainy, and cloudy days. For example, if A > B > C, the environment type corresponding to probability value A is determined as the target object's environment type. For another example, if A > B > C, where the difference between A and B is very small, the environment types corresponding to A and B can be determined as the target object's environment type. Furthermore, the similarity between the multiple environment types is greater than a second preset threshold. In this way, the accuracy of identifying the target object's environment type can be improved.

[0067] For example, the algorithm of the integrated model can use a machine learning regression algorithm. In addition, the models that can be used for the integrated model include but are not limited to logistic regression, random forest, gradient boosting decision tree (Gradient Boosting Decision Tree, GBDT), extreme gradient boosting (eXtreme Gradient Boosting, xgboot), light gradient boosting mechanism (Light Gradient Boosting Machine, LightGBM) and other models, or a deep neural network regression model can be used to perform classification training on the model. In this case, the images obtained by integrating multiple layers and multiple cameras are more accurate in judging the environment type or scene type.

[0068] In general, the image processing method provided by the embodiment of the present disclosure collects multiple images including the target object, each image carries environmental information, and the environmental information includes the category information of the target object, the confidence of the target object, the color category of the target object, and the color ratio; the confidence is used to characterize the credibility of the environmental information of the surrounding environment of the target object; the color ratio is used to characterize the area ratio of the color category of the target object in the entire image; and based on the category information of the target object, the confidence of the target object, the color category of the target object, and the color ratio, the one or more environmental types in which the target object is located in at least one image are determined. In this way, the recognition accuracy of the type of environment in which the target object is located can be improved.

[0069] like Figure 4 FIG. 1 is another image processing method provided in an embodiment of the present application. The method is intended to illustrate that whether the surrounding environment of a target object changes over time can be determined based on time information associated with different images of the target object. The method may include but is not limited to the following steps:

[0070] Step S401: Obtain the color ratio of each color in the first image and the second image respectively.

[0071] Step S402: Determine whether the color ratios of each color corresponding to the first image and the second image are consistent. If they are consistent, execute step S403; if not, execute step S404.

[0072] Step S403: Determine whether the surrounding environment of the target object at the second time information has changed.

[0073] Step S404: Determine that the surrounding environment of the target object at the second time information has not changed.

[0074] In this way, a more detailed analysis of urban colors can be performed.

[0075] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this disclosure is not limited by the order of the actions described, because according to this disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required for this disclosure.

[0076] It should be further explained that although Figure 3 、 Figure 4The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 3 、 Figure 4 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0077] Combined with the above Figure 1-Figure 4 The image processing method of the embodiment of the present application is described in detail. In order to facilitate better implementation of the above-mentioned solution of the embodiment of the present application, correspondingly, relevant devices and equipment for cooperating in implementing the above-mentioned solution are also provided below.

[0078] See also Figure 5 , is a schematic structural diagram of an image processing device 50 provided in an embodiment of the present application, which may include:

[0079] The image acquisition unit 500 is configured to acquire at least one image including a target object; each image carries environmental information, including the category information of the target object, the confidence level of the target object, the color category of the target object, and the color ratio; the confidence level is used to represent the credibility of the environmental information of the target object's surrounding environment; and the color ratio is used to represent the area ratio of the target object's color category in the entire image.

[0080] The processing unit 502 is configured to determine one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category and color proportion of the target object.

[0081] In a possible implementation, the processing unit 502 is specifically configured to:

[0082] At least one image of the target object is processed by a trained machine learning model to determine one or more types of environments in which the target object in the at least one image is located, wherein the trained machine learning model is a model trained based on at least one sample image, and each sample image is associated with annotation data, and the annotation data includes object category information, object confidence, object color category, and color proportion.

[0083] In a possible implementation, the processing unit 502 is specifically configured to:

[0084] Processing each of the at least one image separately, and determining a probability distribution of the type of environment in which the target object is located in each image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion;

[0085] The probability distribution corresponding to each of the images is integrated to obtain one or more environment types in which the target object in the at least one image is located.

[0086] In a possible implementation, the imaging quality corresponding to each of the at least one image of the target object is greater than a first preset threshold.

[0087] In a possible implementation, the color proportion of the target object is a main color proportion.

[0088] In a possible implementation, the similarity between the multiple environment types is greater than a second preset threshold.

[0089] In a possible implementation, the at least one image includes a first image and a second image, the first image carries first time information, the second image carries second time information, and the second time information is a time point after the first time information; the apparatus 50 further includes:

[0090] an environment change determining unit 504, configured to obtain a color ratio of each color in the first image and the second image respectively;

[0091] When the color ratios of each color corresponding to the first image and the second image are inconsistent, it is determined that the surrounding environment of the target object at the second time information has changed.

[0092] It should be noted that each device in the above system may also include other units. The specific implementation of each device and unit can refer to the relevant description in the above method embodiment, which will not be repeated here.

[0093] In order to better implement the above-mentioned solution of the embodiment of the present application, the present application also provides an electronic device 60, which is described in detail below with reference to the accompanying drawings:

[0094] like Figure 6The structural diagram of the electronic device provided in the embodiment of the present application is shown, and the electronic device 600 may include a processor 601, a memory 604 and a communication module 605. The processor 601, the memory 604 and the communication module 605 may be interconnected through a bus 606. The memory 604 may be a high-speed random access memory (RAM) memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 604 may optionally be at least one storage system located away from the aforementioned processor 601. The memory 604 is used to store application code, and may include an operating system, a network communication module, a user interface module and a data processing program. The communication module 605 is used to interact with external devices for information; the processor 601 is configured to call the program code and perform the following steps:

[0095] Acquire at least one image including a target object; wherein each image carries environmental information, the environmental information including category information of the target object, a confidence level of the target object, a color category of the target object, and a color ratio; the confidence level is used to represent the credibility of the environmental information of the target object's surrounding environment; and the color ratio is used to represent the area ratio of the target object's color category in the entire image;

[0096] One or more environment types in which the target object is located in the at least one image are determined based on the category information of the target object, the confidence level of the target object, the color category and the color proportion of the target object.

[0097] The processor 601 determines one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion, including:

[0098] At least one image of the target object is processed by a trained machine learning model to determine one or more types of environments in which the target object in the at least one image is located, wherein the trained machine learning model is a model trained based on at least one sample image, and each sample image is associated with annotation data, and the annotation data includes object category information, object confidence, object color category, and color proportion.

[0099] The processor 601 determines one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion, including:

[0100] Processing each of the at least one image separately, and determining a probability distribution of the type of environment in which the target object is located in each image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion;

[0101] The probability distribution corresponding to each of the images is integrated to obtain one or more environment types in which the target object in the at least one image is located.

[0102] Wherein, the imaging quality corresponding to each of the at least one image of the target object is greater than a first preset threshold.

[0103] The color ratio of the target object is the main color ratio.

[0104] Wherein, the similarity between the multiple environment types is greater than a second preset threshold.

[0105] The at least one image includes a first image and a second image, the first image carries first time information, the second image carries second time information, and the second time information is a time point after the first time information; the processor 601 is further configured to:

[0106] respectively obtaining a color ratio of each color in the first image and the second image;

[0107] When the color ratios of each color corresponding to the first image and the second image are inconsistent, it is determined that the surrounding environment of the target object at the second time information has changed.

[0108] The embodiments of the present application also provide a computer storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer or processor, cause the computer or processor to execute one or more steps of the method described in any of the above embodiments. If the various components of the above-mentioned device are implemented in the form of software functional units and sold or used as independent products, they can be stored in the computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, and the computer product is stored in a computer-readable storage medium.

[0109] The computer-readable storage medium may be an internal storage unit of the device described in the aforementioned embodiment, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of the device, such as an equipped plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Furthermore, the computer-readable storage medium may include both an internal storage unit and an external storage device of the device. The computer-readable storage medium is used to store the computer program and other programs and data required by the device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0110] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0111] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.

[0112] The modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.

[0113] It is understood that those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the various embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0114] Those skilled in the art will appreciate that the functions described in the various illustrative logic blocks, modules, and algorithm steps disclosed in the various embodiments of this application can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions described in the various illustrative logic blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to tangible media, such as data storage media, or communication media including any media that facilitates the transfer of computer programs from one place to another (e.g., according to a communication protocol). In this way, computer-readable media can generally correspond to (1) non-transitory tangible computer-readable storage media, or (2) communication media, such as signals or carrier waves. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described in this application. A computer program product can include computer-readable media.

[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0116] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0118] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0119] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0120] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An image processing method, characterized in that: include: Acquire at least one image including a target object; wherein each image carries environmental information, the environmental information including category information of the target object, a confidence level of the target object, a color category of the target object, and a color ratio; the confidence level is used to represent the credibility of the environmental information of the target object's surrounding environment; and the color ratio is used to represent the area ratio of the target object's color category in the entire image; Determining one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category and the color proportion of the target object; The determining, based on the category information of the target object, the confidence level of the target object, the color category and the color proportion of the target object, one or more types of environments in which the target object is located in the at least one image includes: At least one image of the target object is processed by a trained machine learning model to determine one or more types of environments in which the target object in the at least one image is located, wherein the trained machine learning model is a model trained based on at least one sample image, and each sample image is associated with annotation data, and the annotation data includes object category information, object confidence, object color category, and color proportion.

2. The method according to claim 1, wherein The determining, based on the category information of the target object, the confidence level of the target object, the color category and the color proportion of the target object, one or more types of environments in which the target object is located in the at least one image includes: Processing each of the at least one image separately, and determining a probability distribution of the type of environment in which the target object is located in each image based on the category information of the target object, the confidence level of the target object, the color category of the target object, and the color proportion; The probability distribution corresponding to each of the images is integrated to obtain one or more environment types in which the target object in the at least one image is located.

3. The method according to any one of claims 1 to 2, characterized in that The imaging quality corresponding to each of the at least one image of the target object is greater than a first preset threshold.

4. The method according to any one of claims 1 to 2, wherein: The color proportion of the target object is the main color proportion.

5. The method according to any one of claims 1 to 2, characterized in that The similarity between the multiple environment types is greater than a second preset threshold.

6. The method according to claim 1, wherein The at least one image includes a first image and a second image, the first image carries first time information, the second image carries second time information, and the second time information is a time point after the first time information; the method further includes: respectively obtaining a color ratio of each color in the first image and the second image; When the color ratios of each color corresponding to the first image and the second image are inconsistent, it is determined that the surrounding environment of the target object at the second time information has changed.

7. An image processing device, characterized in that: include: An image acquisition unit, configured to acquire at least one image including a target object; wherein each image carries environmental information, including category information of the target object, a confidence level of the target object, a color category of the target object, and a color ratio; the confidence level is used to represent the credibility of the environmental information of the target object's surrounding environment; and the color ratio is used to represent the area ratio of the target object's color category in the entire image; a processing unit, configured to determine one or more types of environments in which the target object is located in the at least one image based on the category information of the target object, the confidence level of the target object, the color category and color proportion of the target object; The processing unit is specifically used to: At least one image of the target object is processed by a trained machine learning model to determine one or more types of environments in which the target object in the at least one image is located, wherein the trained machine learning model is a model trained based on at least one sample image, and each sample image is associated with annotation data, and the annotation data includes object category information, object confidence, object color category, and color proportion.

8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for adjusting color tendency of object in image and mobile terminal

    CN105760868A

  • Mountain farmland environment identification method and device, electronic equipment and storage medium

    CN114694056A

  • Image processing method and device, equipment and storage medium

    CN114998716A