Model processing method and apparatus, storage medium, and computer device

By optimizing the image processing model by constraining the distance between image sample pairs, constructing a loss function and setting a distance interval, the problem of low efficiency in improving the performance of existing image processing models is solved, and efficient and accurate image processing is achieved.

CN114511872BActive Publication Date: 2026-03-27ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, image processing models have low efficiency in improving performance and require a large number of pre-labeled samples, resulting in a complex workload.

Method used

The image processing model is optimized by constraining a first distance between positive sample pairs and a second distance between negative sample pairs in the image samples. A loss function is constructed to minimize the first distance and a predetermined distance interval is set to optimize the image processing model.

Benefits of technology

It achieves efficient optimization of image processing models, improves the accuracy and efficiency of image processing, and solves the problem of low efficiency when improving the performance of image processing models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511872B_ABST
    Figure CN114511872B_ABST
Patent Text Reader

Abstract

The application discloses a model processing method and device, a storage medium and computer equipment. The method comprises the following steps: acquiring a first pedestrian local area image in a first pedestrian image and a second pedestrian image; and processing the first pedestrian local area image and the second pedestrian image by using an image processing model to obtain a second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is a distance between positive sample pairs in an image sample, and the second distance is a distance between negative sample pairs in the image sample. The application solves the technical problem of low efficiency in improving the performance of the image processing model in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer, in particular to a model processing method and device, a storage medium and a computer device. BACKGROUND

[0002] Deep learning refers to a collection of algorithms for solving various problems such as images and texts on multi-layer neural networks. Image processing technology is an important application direction of deep learning technology, and the application scenarios are extremely wide. In the related technology, the image processing technology generally adopts an image processing model obtained by training a large number of image samples. If the image processing model needs to obtain a more accurate processing rate, the number of samples that need to be pre-labeled is extremely large, and the workload is complex. Therefore, in the related technology, there is a problem of low efficiency in improving the performance of the image processing model.

[0003] At present, no effective solution has been proposed for the above problems. SUMMARY

[0004] The embodiments of the present application provide a model processing method and device, a storage medium and a computer device to at least solve the technical problem of low efficiency in improving the performance of the image processing model in the related technology.

[0005] According to an aspect of the embodiments of the present application, a model processing method is provided, comprising: obtaining a first pedestrian local area image in a first pedestrian image and a second pedestrian image; processing the first pedestrian local area image and the second pedestrian image by using an image processing model to obtain a second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is the distance between positive sample pairs in the image sample, and the second distance is the distance between negative sample pairs in the image sample.

[0006] Optionally, the first distance is at least a predetermined distance interval smaller than the second distance.

[0007] Optionally, the processing the first pedestrian local region image and the second pedestrian image by using the image processing model to obtain the second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image comprises: obtaining a size of the second pedestrian image; performing alignment operation on the first pedestrian local region image by using an alignment module in the image processing model to obtain the first pedestrian local region image with the same size as the second pedestrian image; extracting first pedestrian features of the first pedestrian local region image with the same size as the second pedestrian image and extracting second pedestrian features of the second pedestrian image by using a feature extraction module in the image processing model; and processing the first pedestrian features and the second pedestrian features by using a processing module in the image processing model to obtain the second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image.

[0008] According to an aspect of an embodiment of the present application, a model processing method is provided, comprising: processing a first image by using an image processing model to obtain a second image corresponding to the first image, wherein the first image and the second image are a positive sample pair; obtaining a third image with a similarity lower than a predetermined threshold to the second image, wherein the first image and the third image are a negative sample pair; determining a first distance between the positive sample pair and a second distance between the negative sample pair; and optimizing the image processing model according to the first distance and the second distance.

[0009] Optionally, the optimizing the image processing model according to the first distance and the second distance comprises: constructing a loss function of the image processing model according to the first distance and the second distance, wherein the loss function requires the first distance to be smaller than the second distance; and optimizing the image processing model by minimizing the loss function.

[0010] Optionally, the constructing the loss function of the image processing model according to the first distance and the second distance comprises: determining a predetermined distance interval; and constructing the loss function of the image processing model according to the first distance, the second distance and the predetermined distance interval, wherein the loss function requires the first distance to be at least the predetermined distance interval smaller than the second distance.

[0011] Optionally, the constructing the loss function of the image processing model according to the first distance, the second distance and the predetermined distance interval comprises: constructing the loss function of the image processing model by the following formula:

[0012]

[0013] Wherein, L is a loss value of the loss function, B is a number of positive sample pairs and negative sample pairs, a is the first image, p is the second image, n is the third image, is a first distance between the ith positive sample pair, is a second distance between the ith negative sample pair, and a is the predetermined distance interval.

[0014] Optionally, the first image is an image of a first local region in a to-be-processed image, the second image is an image of a local region in a target image, and the third image is an image of a second local region in the target image.

[0015] Optionally, processing the first image by using the image processing model to obtain the second image corresponding to the first image comprises: obtaining a size of the target image; performing an alignment operation on the first image by using an alignment module in the image processing model to obtain a first image with the same size as the target image; extracting a first feature of the first image with the same size as the target image and extracting a second feature of the target image by using a feature extraction module in the image processing model; and processing the first feature and the second feature by using a processing module in the image processing model to obtain the second image corresponding to the first image.

[0016] Optionally, obtaining the third image with a similarity to the second image lower than a predetermined threshold comprises: obtaining a plurality of second local regions, wherein an intersection of the plurality of second local regions and the first local region is less than the predetermined threshold; selecting a second local region from the plurality of second local regions, and taking an image of the selected second local region as the third image.

[0017] Optionally, the method further comprises: obtaining an image of a local region in a fourth image and a fifth image; and processing the image of the local region in the fourth image and the fifth image by using the optimized image processing model to obtain a corresponding region image in the fifth image corresponding to the local region in the fourth image.

[0018] Optionally, the image processing model comprises an image recognition model, which is configured to recognize at least one of the following predetermined local regions: a local region in a pedestrian image including a person, a local region in an object image including an object, and a local region in a scene image including a person and an object.

[0019] Optionally, the predetermined local region comprises a local region of an image frame intercepted from a monitoring image video.

[0020] According to another aspect of the embodiments of the present application, a model processing method is further provided, comprising: receiving a model optimization instruction through an interactive interface; receiving a first image through the interactive interface based on the model optimization instruction; and displaying a model optimization result on the interactive interface, wherein the model optimization result comprises an optimized image processing model, and the optimized image processing model is obtained by optimizing an image processing model according to a first distance and a second distance, the first distance is a distance between positive sample pairs, and the second distance is a distance between negative sample pairs, the first image and a second image are a positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, and the first image and a third image are a negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold.

[0021] According to still another aspect of the embodiments of the present application, a model processing method is further provided, comprising: receiving a first region image through a front-end client, wherein the first region image is an image of a local region of a first predetermined image; sending the first region image to a back-end server through the front-end client, and receiving a region processing result returned by the back-end server, wherein the region processing result is obtained by processing using an optimized image processing model, and the region processing result comprises an image of a corresponding region of the first predetermined image in a second predetermined image, wherein the optimized image processing model is obtained by optimizing an image processing model according to a first distance and a second distance, the first distance is a distance between positive sample pairs, and the second distance is a distance between negative sample pairs, the first image and a second image are a positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, the first image and a third image are a negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold; and displaying the region processing result through the front-end client.

[0022] According to an aspect of the embodiments of the present application, a model processing device is further provided, comprising: a first acquisition module, configured to acquire a first pedestrian local region image in a first pedestrian image and a second pedestrian image; and a first processing module, configured to process the first pedestrian local region image and the second pedestrian image using an image processing model to obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is a distance between positive sample pairs in an image sample, and the second distance is a distance between negative sample pairs in the image sample.

[0023] According to an aspect of an embodiment of the present application, a model processing apparatus is also provided, which comprises: a second processing module, configured to process a first image by using the image processing model to obtain a second image corresponding to the first image, wherein the first image and the second image are a positive sample pair; a second acquisition module, configured to acquire a third image with a similarity lower than a predetermined threshold to the second image, wherein the first image and the third image are a negative sample pair; a determination module, configured to determine a first distance between the positive sample pair and a second distance between the negative sample pair; and an optimization module, configured to optimize the image processing model according to the first distance and the second distance.

[0024] According to another aspect of an embodiment of the present application, a model processing apparatus is also provided, which comprises: a first receiving module, configured to receive a model optimization instruction through an interactive interface; a second receiving module, configured to receive a first image through the interactive interface based on the model optimization instruction; and a first display module, configured to display a model optimization result on the interactive interface, wherein the model optimization result comprises an optimized image processing model, and the optimized image processing model is obtained by optimizing an image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image by using the image processing model before optimization, the first image and a third image are the negative sample pair, and the third image is an image with a similarity lower than a predetermined threshold to the second image.

[0025] According to still another aspect of an embodiment of the present application, a model processing apparatus is also provided, which comprises: a third receiving module, configured to receive a first region image through a front-end client, wherein the first region image is an image of a local region of a first predetermined image; a fourth receiving module, configured to send the first region image to a back-end server through the front-end client and receive a region processing result returned by the back-end server, wherein the region processing result is obtained by processing by an optimized image processing model, and the region processing result comprises an image of a corresponding region of the first predetermined image in a second predetermined image, wherein the optimized image processing model is obtained by optimizing an image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image by using the image processing model before optimization, the first image and a third image are the negative sample pair, and the third image is an image with a similarity lower than a predetermined threshold to the second image; and a second display module, configured to display the region processing result on the front-end client.

[0026] In the embodiment of the present application, the first image is processed by using the image processing model to obtain the second image corresponding to the first image, and the positive sample pair is constructed by the first image and the second image; the negative sample pair is constructed by the third image with the similarity lower than the predetermined threshold to the second image and the first image, and the first distance between the positive sample pairs and the second distance between the negative sample pairs are compared to achieve the purpose of constraining the distance between the positive sample pairs and the distance between the negative sample pairs, thereby realizing the technical effect of efficiently optimizing the image processing model, and further solving the technical problem of low efficiency in improving the performance of the image processing model in the related art. BRIEF DESCRIPTION OF DRAWINGS

[0027] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:

[0028] Figure 1 A hardware structure block diagram of a computer terminal for implementing the model processing method is shown;

[0029] Figure 2 is a flowchart of the model processing method one according to the embodiment 1 of the present application;

[0030] Figure 3 is a flowchart of the model processing method two according to the embodiment 1 of the present application;

[0031] Figure 4 is a flowchart of the model processing method three according to the embodiment 1 of the present application;

[0032] Figure 5 is a flowchart of the model processing method four according to the embodiment 1 of the present application;

[0033] Figure 6 is a complete pedestrian arbitrary partial recognition structure schematic diagram according to the optional embodiment of the present application;

[0034] Figure 7 is a schematic diagram of the optimization process of the optimization module according to the optional embodiment of the present application;

[0035] Figure 8 is a structure block diagram of the model processing device one according to the embodiment 2 of the present application;

[0036] Figure 9 is a structure block diagram of the model processing device two according to the embodiment 2 of the present application;

[0037] Figure 10 is a structure block diagram of the model processing device three according to the embodiment 2 of the present application;

[0038] Figure 11 is a structural block diagram of the model processing device four according to the embodiment 2 of the present application;

[0039] Figure 12 is a structural block diagram of a computer terminal according to the embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the embodiment of the present application will be described clearly and completely in the following with reference to the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, not all. Based on the embodiment in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0041] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0042] First, some nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:

[0043] Deep learning: deep learning refers to a collection of algorithms for solving various problems such as images and texts on multi-layer neural networks. Deep learning can be classified into neural networks in general, but there are many variations in specific implementation. The core of deep learning is feature learning, which aims to obtain hierarchical feature information through hierarchical network to solve the important problem of manual design of features.

[0044] Artificial neural network: artificial neural network is a research hotspot in the field of artificial intelligence since the 1980s. It is an operation model abstracted from the information processing point of view of human brain neural network, and composed of different networks according to different connection modes.

[0045] Person Recognition: In embodiments of the present application, person recognition refers to taking a person image as input, and outputting a high-dimensional feature vector of the person or a person attribute discrimination result through machine learning or the like.

[0046] Partial Person Recognition: In embodiments of the present application, partial person recognition refers to taking a partial person image of an arbitrary region as input, and outputting a high-dimensional feature vector of the person through machine learning or the like. In a picture search application, a similarity comparison result is obtained by calculating the Euclidean distance or cosine distance between high-dimensional feature vectors. Generally speaking, the closer the distance between the feature vectors, the higher the semantic similarity of the persons corresponding to the feature vectors, that is, the more likely they are the same person.

[0047] Correspondence Learning: In embodiments of the present application, correspondence learning refers to an algorithm automatically finding a semantically corresponding local region in a target image according to a given local region without manual annotation. For example, the algorithm can find the coordinate position of a foot region in another person image according to a human foot region input.

[0048] Embodiment 1

[0049] According to embodiments of the present application, a method embodiment of a model processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0050] The method embodiment provided in Embodiment 1 of the present application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a model processing method is shown. As shown in Figure 1 the computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include but is not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 104 for storing data, and a transmission device for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1The illustrated architecture is merely an example and does not impose a limitation on the architecture of the electronic device. For example, the computer terminal 10 can further include more or less components than those shown, or have a different configuration of components than those shown. Figure 1 Figure 1 The illustrated architecture is merely an example and does not impose a limitation on the architecture of the electronic device. For example, the computer terminal 10 can further include more or less components than those shown, or have a different configuration of components than those shown.

[0051] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. Furthermore, the data processing circuitry can be a single independent processing module or incorporated in whole or in part within any of the other elements of the computer terminal 10 (or mobile device). As referred to in embodiments of the present application, the data processing circuitry functions as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.

[0052] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the model processing method of embodiments of the present application. The processor 102 can execute various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implement the vulnerability detection method of the application program described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory remotely disposed relative to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network can include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0053] The transmission device is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device can be a radio frequency (RF) module used to communicate with the Internet in a wireless manner.

[0054] The display can be, for example, a touch screen type liquid crystal display (LCD) that can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0055] In the above operating environment, the present application provides a model processing method as shown in Figure 2 the accompanying drawings.​Figure 2 is a flowchart of the model processing method one according to the embodiment 1 of the present application, as shown in the figure, the method comprises the following steps: Figure 2

[0056] In step S202, a first pedestrian local region image in a first pedestrian image and a second pedestrian image are acquired.

[0057] In step S204, the first pedestrian local region image and the second pedestrian image are processed by using an image processing model to obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, wherein the image processing model is obtained by optimizing the first distance and the second distance, the first distance is the distance between positive sample pairs in the image sample, and the second distance is the distance between negative sample pairs in the image sample.

[0058] Through the above steps, the first pedestrian local region image and the second pedestrian image are processed by using the image processing model to obtain the second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, so as to achieve the purpose of processing the pedestrian local region according to the optimized image processing model. Since the distance between the positive sample pairs and the distance between the negative sample pairs are constrained, the technical effect of efficiently optimizing the image processing model is achieved, and the pedestrian local region can be accurately and efficiently processed.

[0059] As an optional embodiment, the execution subject of the above method can be any computing device capable of performing calculation, for example, it can be a standalone computer terminal, or a computer cluster with strong computing power, or a server deployed with a computing task.

[0060] As an optional embodiment, the first pedestrian image can be various images including a person, can be an image frame in a video stream including a pedestrian obtained by monitoring the environment in a predetermined area, can be a photo obtained by taking a picture of a predetermined area or environment, or can be a picture uploaded to the Internet or downloaded from the Internet through a predetermined channel.

[0061] As an optional embodiment, the first pedestrian local region image can be an image of each part of the pedestrian. The division of the each part can be various, for example, a wide range of division can be performed: upper part and lower part of the body; a more detailed division can be performed: head, face, elbow, foot, etc.; a more detailed division can be performed: fingers, palms, knuckles, etc. The specific division method can be selected flexibly according to specific needs.

[0062] ​As an optional embodiment, the first distance and the second distance are used to identify the accuracy of the identification processing result relative to the object to be processed. For example, if it is a positive sample pair, the smaller the distance, the more accurate the identification processing result, otherwise the less accurate the identification processing result. If it is a negative sample pair, the larger the distance, the more accurate the identification processing result, otherwise the less accurate the identification processing result. The first distance represents the distance between the positive sample pairs, and the smaller the distance, the more accurate the processing result of the image processing model. The second distance represents the distance between the negative sample pairs, and the larger the distance, the more accurate the processing result of the image processing model. If the first distance is smaller than the second distance, it indicates that the image processing model is trained in the direction of optimization, i.e. it becomes more and more optimized. In order to make the optimization of the image processing model reach a certain degree, a distance interval can be set, for example, the first distance is at least a predetermined distance interval smaller than the second distance.

[0063] As an optional embodiment, when the image processing model is used to process the first pedestrian local area image and the second pedestrian image to obtain the second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image, the following method can be used: obtaining the size of the second pedestrian image; using the alignment module in the image processing model to perform an alignment operation on the first pedestrian local area image to obtain a first pedestrian local area image with the same size as the second pedestrian image; using the feature extraction module in the image processing model to extract first pedestrian features of the first pedestrian local area image with the same size as the second pedestrian image, and to extract second pedestrian features of the second pedestrian image; using the processing module in the image processing model to process the first pedestrian features and the second pedestrian features to obtain the second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image. That is, the processing of the image processing model on the first pedestrian local area image and the second pedestrian image can use the model structure of the above alignment module, feature extraction module and processing module to realize the entire processing process. Of course, the above model structure is only an example, and other similar or obvious variations of the above model structure also belong to the present application.

[0064] In the above operating environment, the present application provides a model processing method as shown in Figure 3 . Figure 3 is a flowchart of the model processing method two according to embodiment 1 of the present application, as shown in Figure 3 , the method comprises the following steps:

[0065] Step S302, using an image processing model to process a first image to obtain a second image corresponding to the first image, wherein the first image and the second image are a positive sample pair;

[0066] Step S304: Obtain a third image whose similarity to the second image is lower than a predetermined threshold, wherein the first image and the third image are negative sample pairs;

[0067] Step S306: Determine the first distance between positive sample pairs and the second distance between negative sample pairs;

[0068] Step S308: Optimize the image processing model based on the first distance and the second distance.

[0069] Through the above steps, the first image is processed using an image processing model to obtain a second image corresponding to the first image, and a positive sample pair is constructed from the first image and the second image. A negative sample pair is constructed from a third image whose similarity to the second image is lower than a predetermined threshold and the first image. By comparing the first distance between positive sample pairs and the second distance between negative sample pairs, the distance between positive sample pairs and the distance between negative sample pairs are constrained. This achieves the technical effect of efficiently optimizing the image processing model, and solves the technical problem of low efficiency when improving the performance of image processing models in related technologies.

[0070] As an optional embodiment, the execution subject of the above method can be any computing device that can be used to perform calculations, such as an independent computer terminal, a computer cluster with strong computing power, or a server with computing tasks deployed.

[0071] As an optional embodiment, the image processing model can be optimized based on the first distance and the second distance in various ways. For example, it can be done by constructing a loss function for the image processing model based on the first distance and the second distance, where the loss function requires the first distance to be less than the second distance; and optimizing the image processing model by minimizing the loss function. It should be noted that the forms in which the first distance and the second distance are described can be diverse. For example, they can be Euclidean distances representing the coordinates between the first and second images, or between the first and third images; they can also be cosine distances representing the coordinates between the first and second images, or between the first and third images, etc.; there is no specific limitation here. Furthermore, the form of the loss function constructed here can also be diverse, as long as it can measure the first distance between the first and second images, or the second distance between the first and third images. For example, it can be a triplet loss function, a cross-entropy loss function, etc.

[0072] As an optional embodiment, when the loss function of the image processing model is constructed according to the first distance and the second distance, and the image processing model is optimized by minimizing the loss function, the constructed loss function can be a single type of loss function, such as the above-mentioned triplet loss function or cross-entropy loss function, or a combination of multiple types of loss functions, such as a total loss function obtained by weighted summation of the triplet loss function and the cross-entropy loss function, wherein the weights of each loss function can be determined according to the optimization performance of the corresponding loss function. For example, if the historical optimization performance of the loss function is relatively high, the weight coefficient of its combination can be set to be relatively large, and if the historical optimization performance of the loss function is relatively low, the weight coefficient of its combination can be set to be relatively small.

[0073] As an optional embodiment, the above-mentioned loss function requires the first distance to be less than the second distance, that is, the distance between the positive sample pairs is required to be less than the distance between the negative sample pairs. If the second distance is greater than or equal to the first distance, the loss value obtained by the above-mentioned loss function will be larger, and the optimization effect of the image processing model will be poor. Therefore, by requiring the first distance to be less than the second distance, the image processing model can be optimized more and more.

[0074] As an optional embodiment, the above-mentioned loss function requires the first distance to be less than the second distance, that is, the distance between the positive sample pairs is required to be less than the distance between the negative sample pairs. If the second distance is greater than or equal to the first distance, the loss value obtained by the above-mentioned loss function will be larger, and the optimization effect of the image processing model will be poor. Therefore, by requiring the first distance to be less than the second distance, the image processing model can be optimized more and more.

[0075] As an optional embodiment, according to the first distance, the second distance and the predetermined distance interval, the loss function of the image processing model can be constructed in the following manner, for example, the loss function of the image processing model can be constructed by the following formula:

[0076]

[0077] wherein L is the loss value of the loss function, B is the number of positive sample pairs and negative sample pairs, a is the first image, p is the second image, and n is the third image. is a first distance between the i-th pair of positive samples, is a second distance between the i-th pair of negative samples, and a is a predetermined distance interval.

[0078] As an optional embodiment, the image processing model can be used to process a local region in a complete image. That is, for a local region in a complete image, a corresponding region corresponding to the local region is processed from a target image. For example, the first image can be an image of a first local region in a to-be-processed image, the second image can be an image of a local region in a target image, and the third image can be an image of a second local region in the target image. Therefore, the second image is a corresponding region image in the target image corresponding to the first local region in the to-be-processed image, that is, the first image and the second image are positive sample images. The third image can be an image of a region in the target image that does not correspond to the first local region in the to-be-processed image, that is, the first image and the third image are negative sample images.

[0079] As an optional embodiment, the image processing model can be used to process a local region in a complete image. That is, for a local region in a complete image, a corresponding region corresponding to the local region is processed from a target image. For example, the first image can be an image of a first local region in a to-be-processed image, the second image can be an image of a local region in a target image, and the third image can be an image of a second local region in the target image. Therefore, the second image is a corresponding region image in the target image corresponding to the first local region in the to-be-processed image, that is, the first image and the second image are positive sample images. The third image can be an image of a region in the target image that does not correspond to the first local region in the to-be-processed image, that is, the first image and the third image are negative sample images.

[0080] As an optional embodiment, when the third image with a similarity lower than the predetermined threshold to the second image is obtained, the following method can be used: first, a plurality of second local regions are obtained, wherein the intersection of each of the plurality of second local regions and the first local region is less than the predetermined threshold; then, a second local region is selected from the plurality of second local regions, and the image of the selected second local region is taken as the third image. The intersection of each of the plurality of second local regions and the first local region is less than the predetermined threshold, that is, the similarity of each of the plurality of second local regions and the first local region is less than the predetermined threshold. The predetermined threshold can be flexibly determined according to specific requirements. For example, the predetermined threshold can be a specific numerical value, for example, the number of pixels of the intersection of each of the plurality of second local regions and the first local region is less than 300,000 pixels, or a relative ratio value, for example, the intersection of each of the plurality of second local regions and the first local region is less than 30%.

[0081] As an optional embodiment, after obtaining the optimized image processing model, the image can be processed by using the optimized image processing model, for example, a corresponding local image corresponding to the image of the predetermined local region can be processed in the predetermined image. Using the optimized image processing model for processing effectively improves the accuracy of image processing. For example, the image can be processed by using the optimized image processing model in the following manner: obtaining the image of the local region in the fourth image and the fifth image; processing the image of the local region in the fourth image and the fifth image by using the optimized image processing model, to obtain the corresponding region image in the fifth image corresponding to the local region in the fourth image.

[0082] As an optional embodiment, the image processing model described above includes an image recognition model, which can be used for recognition of local regions of various types of images, for example, can be used for recognition of local regions in a pedestrian image including a person; for example, can be used for recognition of local regions in a scene image including a person and an object; for example, can be used for recognition of local regions in an object image including an object.

[0083] As an optional embodiment, the pedestrian image described above can also include various types, for example, can include image frames intercepted from a monitoring image video. Through recognition of the image frames intercepted from the monitoring image video, the local region of the person included in the image frames can be efficiently recognized, thereby improving the effectiveness of monitoring.

[0084] The application also provides a model processing method as shown in Figure 4 . Figure 4 is a flowchart of the model processing method three according to Embodiment 1 of the application. As shown in Figure 4 , the method comprises the following steps:

[0085] Step S402, receiving a model optimization instruction through an interactive interface;

[0086] Step S404, receiving a first image through the interactive interface based on the model optimization instruction;

[0087] Step S406, displaying a model optimization result on the interactive interface, wherein the model optimization result includes an optimized image processing model, and the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, the first image and a third image are the negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold.

[0088] Through the above steps, the model optimization instruction is received through the interactive interface, and the model optimization result is displayed on the interactive interface, wherein in the process of obtaining the optimized image processing model, the first image is processed using the image processing model to obtain the second image corresponding to the first image, and the positive sample pair is constructed by the first image and the second image; the negative sample pair is constructed by the third image with a similarity to the second image lower than the predetermined threshold and the first image; by comparing the first distance between the positive sample pair and the second distance between the negative sample pair, the distance between the positive sample pair and the distance between the negative sample pair are constrained, thereby achieving the technical effect of efficiently optimizing the image processing model, and further solving the technical problem of low efficiency in improving the performance of the image processing model in the related art.

[0089] The application also provides a model processing method as shown in Figure 5 . Figure 5 is a flowchart of the model processing method four according to Embodiment 1 of the application. As shown in Figure 5 , the method includes the following steps:

[0090] Step S502, receiving a first region image through a front-end client, wherein the first region image is an image of a local region of a first predetermined image;

[0091] In step S504, the front-end client sends the first region image to the back-end server, and receives a region processing result returned by the back-end server, wherein the region processing result is processed by the optimized image processing model, and the region processing result includes an image of a corresponding region in the second predetermined image corresponding to the local region of the first predetermined image, wherein the optimized image processing model is obtained by optimizing the image processing model according to the first distance and the second distance, the first distance is the distance between the positive sample pairs, the second distance is the distance between the negative sample pairs, the first image and the second image are the positive sample pairs, the second image is an image corresponding to the first image obtained by processing the first image by using the image processing model before optimization, and the first image and the third image are the negative sample pairs, and the third image is an image with a similarity to the second image lower than a predetermined threshold.

[0092] In step S506, the front-end client displays the region processing result.

[0093] Through the above steps, after the front-end client receives the image of the local region, the image is sent to the back-end server, and the corresponding region processing result is returned by the back-end server, wherein the region processing result is processed by the optimized image processing model, in the process of optimizing the image processing model, the first image is processed by using the image processing model to obtain the second image corresponding to the first image, and the positive sample pairs are constructed by using the first image and the second image; the negative sample pairs are constructed by using the third image with a similarity to the second image lower than a predetermined threshold and the first image, and the first distance between the positive sample pairs and the second distance between the negative sample pairs are compared, so as to constrain the distance between the positive sample pairs and the distance between the negative sample pairs, thereby realizing the technical effect of efficiently optimizing the image processing model, and solving the technical problem of low efficiency in improving the performance of the image processing model in the related art.

[0094] In the present application, an optional embodiment is also provided, which is more specific. In the optional embodiment, the image processing model is an image recognition model, the to-be-recognized image is taken as a pedestrian image, and the basis for recognition is the image of the local region of the pedestrian image, and the object to be recognized is the corresponding region corresponding to the local region of the pedestrian image in the target image.

[0095] The pedestrian recognition technology is an important application direction of the image recognition technology, but most of the current pedestrian recognition technologies assume that the input is a complete pedestrian image, and incomplete input is rarely considered, and the incomplete input will have a great negative impact on the recognition performance.

[0096] The main technical points of the pedestrian arbitrary local recognition generally include two aspects. The first aspect is to learn a multi-size feature to adapt to an arbitrary input size. The second aspect is to find a region in a target image related to a given local input. A function such as least square can be used to model the relationship between the given local region feature and the target image feature, so as to obtain the weight coefficients between different regions of the target image feature. When performing feature similarity comparison, the given local region feature and the weighted target image feature are used for calculation. However, the method of modeling the relationship by using the least square function has low efficiency and high complexity, and is not flexible, because the semantic distance between regions is modeled by using a human-set function.

[0097] Therefore, in the optional embodiment of the present application, a dual constraint technology based on a local corresponding region is proposed. The dual constraint technology uses the duality of the ability to find a corresponding region according to a given input, that is, if the algorithm can find a corresponding local region y in a pedestrian image Y according to a local region x in a given pedestrian image X, it can also output the local region x in the pedestrian image X according to the found corresponding region y in the reverse direction.

[0098] Figure 6 is a complete pedestrian arbitrary local recognition structure schematic diagram according to the optional embodiment of the present application, as shown in Figure 6 x p is an input arbitrary local region image, y is a target image, x p is aligned through a spatial transform network (Spatial Transform Networks) based alignment module R, and an image with y is obtained through the same feature extraction module F to obtain corresponding features. The dual constraint technology based on a local region proposed in the optional embodiment of the present application is used for an optimization module G. The function of the optimization module G is to output the target region position (the dashed box in the right oblique region) corresponding to the semantic of the input region in the target image feature (the right oblique region) according to the given local region feature (the solid box in the left oblique region) and the target image feature.

[0099] Figure 7 is an optimization process schematic diagram of the optimization module according to the optional embodiment of the present application, as shown in Figure 7 , pixel domain representation is used instead of feature domain representation. In Figure 7In the middle, the given local region is represented by the area in the left solid box, the area in the right dashed box represents the local region on the target image that is semantically corresponding to the local region predicted by the module G, and the area in the right solid box is a randomly sampled local region on the target image (ensuring that the intersection with the area in the right dashed box is less than 30%). In order to make the local corresponding region predicted by the module G better, the area in the left solid box and the area in the right dashed box can be positive sample pairs, the area in the left solid box and the area in the right solid box can be negative sample pairs, and the distance between the features of the positive sample pairs is constrained to be less than the distance between the features of the negative sample pairs. This constraint can be expressed by the formula:

[0100]

[0101] where B is the batch size, that is, the number of positive sample pairs and negative sample pairs mentioned above, a and p are positive sample pairs, a and n are negative sample pairs, and a is a hyperparameter that controls the optimization boundary, corresponding to the predetermined distance interval mentioned above.

[0102] It should be noted that in the optional embodiments of the present application, the alignment module and the feature extraction module are not limited to a certain specific form, and any feature extraction structure that can preserve relevant position information is applicable. In the optional embodiments of the present application, the module G is usually a convolutional neural network structure, but is not limited to this structure form, and any feature extraction structure is applicable. In the optional embodiments of the present application, the minimum distance can be achieved by minimizing the Euclidean distance of the coordinates of the two regions (blue blocks and pink blocks), but is not limited to this distance form, and any loss function that measures the distance between the two regions is applicable. In addition, in the optional embodiments of the present application, the negative samples are randomly sampled results, but are not limited to a specific sampling strategy.

[0103] The method provided by the above optional embodiments can effectively find the corresponding region in the target image according to the given local region image without additional manual annotation, thereby enhancing the performance of pedestrian arbitrary local recognition. In addition, the method provided by the optional embodiments of the present application only needs one inference when deployed, ensuring real-time efficiency, and compared with the method used in the related art, the method provided by the present application achieves higher recognition accuracy on public data sets.

[0104] Through the optional embodiment, a pedestrian arbitrary local recognition technology based on a corresponding region triplet constraint is provided. The corresponding region triplet constraint technology can utilize the characteristic that a feature distance between better corresponding regions is less than a feature distance between random corresponding regions, and realize the role of determining better corresponding regions through triplet loss function constraint. The corresponding region triplet constraint technology firstly models local regions into a triplet relationship, and reduces a feature distance of a positive sample pair and increases a feature distance of a negative sample pair by constructing the positive and negative sample pairs. The corresponding region triplet constraint technology can effectively predict a corresponding local region in a target image according to a given local region, and improve pedestrian arbitrary local recognition performance on this basis.

[0105] It should be noted that, for each of the above method embodiments, in order to simply describe, each is expressed as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0106] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, and of course it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the method of each embodiment of the present application.

[0107] Embodiment 2

[0108] According to the embodiments of the present application, a device for implementing the above model processing method is also provided, Figure 8 is a structural block diagram of the model processing device one according to the embodiment 2 of the present application, as Figure 8 shown, the device includes a first acquisition module 82 and a first processing module 84, and the device will be described below.

[0109] The first obtaining module 82 is configured to obtain a first pedestrian local region image in a first pedestrian image and a second pedestrian image; the first processing module 84, connected to the first obtaining module 82, is configured to process the first pedestrian local region image and the second pedestrian image by using an image processing model to obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance being a distance between positive sample pairs in an image sample, and the second distance being a distance between negative sample pairs in the image sample.

[0110] It should be noted that the first obtaining module 82 and the first processing module 84 correspond to steps S202 to S204 in Embodiment 1, and the modules and the corresponding steps have the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the modules can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0111] According to the embodiments of the present application, a device for implementing the above-mentioned model processing method is further provided, Figure 9 is a structural block diagram of the model processing device two according to Embodiment 2 of the present application, as Figure 9 shown, the device comprises a second processing module 92, a second obtaining module 94, a determination module 96 and an optimization module 98, which are described below.

[0112] The second processing module 92 is configured to process a first image by using an image processing model to obtain a second image corresponding to the first image, wherein the first image and the second image are a positive sample pair; the second obtaining module 94, connected to the second processing module 92, is configured to obtain a third image with a similarity lower than a predetermined threshold to the second image, wherein the first image and the third image are a negative sample pair; the determination module 96, connected to the second obtaining module 94, is configured to determine a first distance between the positive sample pairs and a second distance between the negative sample pairs; and the optimization module 98, connected to the determination module 96, is configured to optimize the image processing model according to the first distance and the second distance.

[0113] It should be noted that the second processing module 92, the second obtaining module 94, the determination module 96 and the optimization module 98 correspond to steps S302 to S308 in Embodiment 1, and the modules and the corresponding steps have the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the modules can be run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0114] According to the embodiments of the present application, a device for implementing the above-mentioned model processing method three is further provided, Figure 10is a structural block diagram of the model processing device three according to the embodiment 2 of the present application, as shown in the figure, the device comprises a first receiving module 102, a second receiving module 104 and a first display module 106, and the device will be described below. Figure 10

[0115] The first receiving module 102 is configured to receive a model optimization instruction through an interactive interface; the second receiving module 104 is connected to the first receiving module 102 and configured to receive a first image through the interactive interface based on the model optimization instruction; and the first display module 106 is connected to the second receiving module 104 and configured to display a model optimization result on the interactive interface, wherein the model optimization result comprises an optimized image processing model, the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a pair of positive samples, the second distance is a distance between a pair of negative samples, the first image and a second image are the pair of positive samples, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, the first image and a third image are the pair of negative samples, and the third image is an image with a similarity to the second image lower than a predetermined threshold.

[0116] It should be noted that the first receiving module 102, the second receiving module 104 and the first display module 106 correspond to steps S402 to S406 in the embodiment 1, and the above modules have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in the above embodiment 1. It should be noted that the above modules can run in the computer terminal 10 provided in the embodiment 1 as a part of the device.

[0117] According to the embodiment of the present application, a device for implementing the above-mentioned model processing method four is further provided, Figure 11 is a structural block diagram of the model processing device four according to the embodiment 2 of the present application, as shown in the figure, the device comprises a third receiving module 112, a fourth receiving module 114 and a second display module 116, and the device will be described below. Figure 11

[0118] ​​The third receiving module 112 is configured to receive the first area image by the front-end client, wherein the first area image is an image of a local area of the first predetermined image; the fourth receiving module 114 is connected to the third receiving module 112 and configured to send the first area image to the back-end server by the front-end client and receive the area processing result returned by the back-end server, wherein the area processing result is obtained by processing the optimized image processing model, and the area processing result includes an image of a corresponding area of the second predetermined image to the local area of the first predetermined image, wherein the optimized image processing model is obtained by optimizing the image processing model according to the first distance and the second distance, the first distance is the distance between the positive sample pairs, the second distance is the distance between the negative sample pairs, the first image and the second image are the positive sample pairs, the second image is the image corresponding to the first image obtained by processing the first image by the image processing model before optimization, and the first image and the third image are the negative sample pairs, and the third image is the image with a similarity lower than a predetermined threshold to the second image; and the second display module 116 is connected to the fourth receiving module 114 and configured to display the area processing result by the front-end client.

[0119] It should be noted that the third receiving module 112, the fourth receiving module 114 and the second display module 116 correspond to steps S502 to S506 in Embodiment 1, and the above modules have the same instances and application scenarios as the corresponding steps, but are not limited to the above-mentioned Embodiment 1. It should be noted that the above modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.

[0120] Embodiment 3

[0121] The embodiments of the present application can provide a computer terminal, which can be any one of the computer terminal devices in the computer terminal group. Alternatively, in the present embodiment, the above computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0122] Alternatively, in the present embodiment, the above computer terminal can be located in at least one of the network devices in the computer network.

[0123] In the present embodiment, the computer terminal can execute the program code of the following steps in the model processing method of the application program: processing the first image by the image processing model to obtain the second image corresponding to the first image, wherein the first image and the second image are the positive sample pairs; obtaining the third image with a similarity lower than a predetermined threshold to the second image, wherein the first image and the third image are the negative sample pairs; determining the first distance between the positive sample pairs and the second distance between the negative sample pairs; and optimizing the image processing model according to the first distance and the second distance.

[0124] Alternatively, Figure 12is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in Figure 12 the computer terminal can include one or more (only one is shown in the figure) processors 122, memory 124, etc.

[0125] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the model processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned model processing method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the computer terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0126] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtaining a first pedestrian local area image in a first pedestrian image and a second pedestrian image; and processing the first pedestrian local area image and the second pedestrian image by using an image processing model to obtain a second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is a distance between positive sample pairs in an image sample, and the second distance is a distance between negative sample pairs in the image sample.

[0127] Optionally, the processor can further execute program codes of the following steps: the first distance is at least smaller than the second distance by a predetermined distance interval.

[0128] Optionally, the processor can further execute program codes of the following steps: processing the first pedestrian local area image and the second pedestrian image by using the image processing model to obtain the second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image includes: obtaining a size of the second pedestrian image; performing an alignment operation on the first pedestrian local area image by using an alignment module in the image processing model to obtain a first pedestrian local area image with the same size as the second pedestrian image; extracting a first pedestrian feature of the first pedestrian local area image with the same size as the second pedestrian image and extracting a second pedestrian feature of the second pedestrian image by using a feature extraction module in the image processing model; and processing the first pedestrian feature and the second pedestrian feature by using a processing module in the image processing model to obtain the second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image.

[0129] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: processing a first image by using an image processing model to obtain a second image corresponding to the first image, wherein the first image and the second image are a positive sample pair; obtaining a third image with a similarity to the second image lower than a predetermined threshold, wherein the first image and the third image are a negative sample pair; determining a first distance between the positive sample pair and a second distance between the negative sample pair; and optimizing the image processing model according to the first distance and the second distance.

[0130] Optionally, the processor can further execute program codes of the following steps: optimizing the image processing model according to the first distance and the second distance includes: constructing a loss function of the image processing model according to the first distance and the second distance, wherein the loss function requires the first distance to be smaller than the second distance; and optimizing the image processing model by minimizing the loss function.

[0131] Optionally, the processor can further execute program codes of the following steps: constructing the loss function of the image processing model according to the first distance and the second distance includes: determining a predetermined distance interval; and constructing the loss function of the image processing model according to the first distance, the second distance and the predetermined distance interval, wherein the loss function requires the first distance to be at least the predetermined distance interval smaller than the second distance.

[0132] Optionally, the processor can further execute program codes of the following steps: constructing the loss function of the image processing model according to the first distance, the second distance and the predetermined distance interval includes: constructing the loss function of the image processing model by the following formula:

[0133]

[0134] wherein L is a loss value of the loss function, B is a number of the positive sample pair and the negative sample pair, a is the first image, p is the second image, n is the third image, is the first distance between the i-th positive sample pair, is the second distance between the i-th negative sample pair, and a is the predetermined distance interval.

[0135] Optionally, the processor can further execute program codes of the following steps: the first image is an image of a first local region in a to-be-processed image, the second image is an image of a local region in a target image, and the third image is an image of a second local region in the target image.

[0136] Optionally, the processor can further execute program codes of the following steps: processing the first image by using the image processing model to obtain the second image corresponding to the first image comprises: obtaining a size of the target image; performing an alignment operation on the first image by using an alignment module in the image processing model to obtain the first image with the same size as the target image; extracting a first feature of the first image with the same size as the target image and a second feature of the target image by using a feature extraction module in the image processing model; and processing the first feature and the second feature by using a processing module in the image processing model to obtain the second image corresponding to the first image.

[0137] Optionally, the processor can further execute program codes of the following steps: obtaining the third image with a similarity lower than a predetermined threshold value with the second image comprises: obtaining a plurality of second local regions, wherein an intersection of the plurality of second local regions and the first local region is less than a predetermined threshold value; selecting a second local region from the plurality of second local regions, and taking an image of the selected second local region as the third image.

[0138] Optionally, the processor can further execute program codes of the following steps: obtaining an image of a local region in the fourth image and a fifth image; and processing the image of the local region in the fourth image and the fifth image by using the optimized image processing model to obtain a corresponding region image in the fifth image corresponding to the local region in the fourth image.

[0139] Optionally, the processor can further execute program codes of the following steps: the image processing model comprises an image recognition model, and is used to recognize at least one of the following predetermined local regions: a local region in a pedestrian image comprising a person, a local region in an object image comprising an object, and a local region in a scene image comprising a person and an object.

[0140] Optionally, the processor can further execute program codes of the following steps: the predetermined local region comprises a local region of an image frame intercepted from a monitoring image video.

[0141] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: receiving a model optimization instruction through an interactive interface; receiving a first image through the interactive interface based on the model optimization instruction; and displaying a model optimization result on the interactive interface, wherein the model optimization result includes an optimized image processing model, and the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, and the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, and the first image and a third image are the negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold.

[0142] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: receiving a first area image through a front-end client, wherein the first area image is an image of a local area of a first predetermined image; sending the first area image to a back-end server through the front-end client, and receiving an area processing result returned by the back-end server, wherein the area processing result is obtained by processing using an optimized image processing model, and the area processing result includes an image of a corresponding area of the first predetermined image in a second predetermined image, wherein the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, and the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, the first image and a third image are the negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold; and displaying the area processing result through the front-end client.

[0143] By adopting the embodiment of the present application, a model processing scheme is provided. The first image is processed by the image processing model to obtain the second image corresponding to the first image, and the positive sample pair is constructed by the first image and the second image. The negative sample pair is constructed by the third image with a similarity to the second image lower than a predetermined threshold and the first image. By comparing the first distance between the positive sample pair and the second distance between the negative sample pair, the distance between the positive sample pair and the distance between the negative sample pair are constrained, thereby realizing the technical effect of efficiently optimizing the image processing model, and further solving the technical problem of low efficiency in improving the performance of the image processing model in the related art.

[0144] Those skilled in the art can understand that, Figure 12The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or other terminal device. Figure 12 This does not limit the structure of the electronic device described above. For example, the computer terminal can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the middle, or have a different configuration from that shown. Figure 12 Figure 12

[0145] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the relevant hardware of the terminal device, and the program can be stored in a computer-readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0146] Embodiment 4

[0147] The embodiments of the present application also provide a storage medium. Optionally, in the present embodiment, the storage medium can be used to save the program code executed by the model processing method provided in Embodiment 1.

[0148] Optionally, in the present embodiment, the storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.

[0149] Optionally, in the present embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a first pedestrian local area image in a first pedestrian image and a second pedestrian image; and processing the first pedestrian local area image and the second pedestrian image by using an image processing model to obtain a second pedestrian local area image corresponding to the first pedestrian local area image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is a distance between positive sample pairs in an image sample, and the second distance is a distance between negative sample pairs in the image sample.

[0150] Optionally, in the present embodiment, the storage medium is further configured to store program code for performing the following step: the first distance is at least a predetermined distance interval smaller than the second distance.

[0151] ​​Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: processing the first pedestrian local region image and the second pedestrian image by using the image processing model to obtain the second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, including: obtaining the size of the second pedestrian image; performing an alignment operation on the first pedestrian local region image by using an alignment module in the image processing model to obtain the first pedestrian local region image with the same size as the second pedestrian image; extracting the first pedestrian feature of the first pedestrian local region image with the same size as the second pedestrian image and the second pedestrian feature of the second pedestrian image by using a feature extraction module in the image processing model; and processing the first pedestrian feature and the second pedestrian feature by using a processing module in the image processing model to obtain the second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image.

[0152] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: processing the first image by using the image processing model to obtain the second image corresponding to the first image, wherein the first image and the second image are a positive sample pair; obtaining a third image with a similarity lower than a predetermined threshold to the second image, wherein the first image and the third image are a negative sample pair; determining a first distance between the positive sample pair and a second distance between the negative sample pair; and optimizing the image processing model according to the first distance and the second distance.

[0153] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: optimizing the image processing model according to the first distance and the second distance includes: constructing a loss function of the image processing model according to the first distance and the second distance, wherein the loss function requires the first distance to be smaller than the second distance; and optimizing the image processing model by minimizing the loss function.

[0154] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: constructing the loss function of the image processing model according to the first distance and the second distance includes: determining a predetermined distance interval; and constructing the loss function of the image processing model according to the first distance, the second distance, and the predetermined distance interval, wherein the loss function requires the first distance to be at least the predetermined distance interval smaller than the second distance.

[0155] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: constructing the loss function of the image processing model according to the first distance, the second distance, and the predetermined distance interval includes: constructing the loss function of the image processing model by the following formula:

[0156]

[0157] Wherein, L is a loss value of the loss function, B is a number of positive sample pairs and negative sample pairs, a is the first image, p is the second image, n is the third image, is a first distance between the ith positive sample pair, is a second distance between the ith negative sample pair, and a is a predetermined distance interval.

[0158] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: the first image is an image of a first local region in the to-be-processed image, the second image is an image of a local region in the target image, and the third image is an image of a second local region in the target image.

[0159] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: processing the first image by using the image processing model to obtain the second image corresponding to the first image includes: obtaining a size of the target image; performing an alignment operation on the first image by using an alignment module in the image processing model to obtain the first image with the same size as the target image; extracting a first feature of the first image with the same size as the target image and extracting a second feature of the target image by using a feature extraction module in the image processing model; and processing the first feature and the second feature by using a processing module in the image processing model to obtain the second image corresponding to the first image.

[0160] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: obtaining the third image with a similarity to the second image lower than a predetermined threshold includes: obtaining a plurality of second local regions, wherein the intersection of the plurality of second local regions and the first local region is less than a predetermined threshold; selecting a second local region from the plurality of second local regions, and taking an image of the selected second local region as the third image.

[0161] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: obtaining an image of a local region in the fourth image, and a fifth image; and processing the image of the local region in the fourth image and the fifth image by using the optimized image processing model to obtain a corresponding region image in the fifth image corresponding to the local region in the fourth image.

[0162] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: the image processing model includes an image recognition model, and is configured to recognize at least one of the following predetermined local regions: a local region in a pedestrian image including a person, a local region in an object image including an object, and a local region in a scene image including a person and an object.

[0163] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: the predetermined local region comprises a local region of an image frame intercepted from the monitoring image video.

[0164] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: receiving a model optimization instruction through the interactive interface; receiving a first image through the interactive interface based on the model optimization instruction; and displaying a model optimization result on the interactive interface, wherein the model optimization result comprises an optimized image processing model, and the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, and the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, and the first image and a third image are the negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold.

[0165] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: receiving a first region image through the front-end client, wherein the first region image is an image of a local region of a first predetermined image; sending the first region image to the back-end server through the front-end client, and receiving a region processing result returned by the back-end server, wherein the region processing result is obtained by processing the first region image using the optimized image processing model, and the region processing result comprises an image of a corresponding region of the second predetermined image corresponding to the local region of the first predetermined image, wherein the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, and the second distance is a distance between a negative sample pair, the first image and a second image are the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, the first image and a third image are the negative sample pair, and the third image is an image with a similarity to the second image lower than a predetermined threshold; and displaying the region processing result through the front-end client.

[0166] The above-mentioned serial numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0167] In the above-mentioned embodiments of the application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0168] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0169] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0170] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0171] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0172] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A model processing method characterized by comprising: The method comprises: obtaining a first pedestrian local region image in a first pedestrian image, and a second pedestrian image, wherein the first pedestrian image is used to represent various images including a person, and the first pedestrian local region image is used to represent images of various parts of the person; obtaining a size of the second pedestrian image; performing an alignment operation on the first pedestrian local region image by using an image processing model to obtain a first pedestrian local region image with the same size as the second pedestrian image, so as to obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is a distance between positive sample pairs in an image sample, and the second distance is a distance between negative sample pairs in the image sample, and the first distance is at least smaller than the second distance by a predetermined distance interval.

2. The method of claim 1, wherein, The method comprises: performing an alignment operation on the first pedestrian local region image by using an image processing model to obtain a first pedestrian local region image with the same size as the second pedestrian image, so as to obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, wherein the method comprises: performing an alignment operation on the first pedestrian local region image by using an alignment module in the image processing model to obtain a first pedestrian local region image with the same size as the second pedestrian image; extracting a first pedestrian feature of the first pedestrian local region image with the same size as the second pedestrian image and extracting a second pedestrian feature of the second pedestrian image by using a feature extraction module in the image processing model; 3. A model processing method characterized by, processing the first pedestrian feature and the second pedestrian feature by using a processing module in the image processing model to obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image. The method comprises: obtaining a size of a target image; performing an alignment operation on a first image to obtain a first image with the same size as the target image, so as to obtain a second image corresponding to the first image, wherein the first image is an image of a first local region in a to-be-processed image, the second image is an image of a local region in the target image, and the first image and the second image are a positive sample pair; obtaining a third image with a similarity lower than a predetermined threshold to the second image, wherein the first image and the third image are a negative sample pair; determining a first distance between the positive sample pairs and a second distance between the negative sample pairs; 4. The method of claim 3, wherein, optimizing the image processing model according to the first distance and the second distance, wherein the image processing model is used to perform the model processing method in any one of claims 1 to 2. The method comprises: constructing a loss function of the image processing model according to the first distance and the second distance, wherein the loss function requires that the first distance is smaller than the second distance; optimizing the image processing model by minimizing the loss function.

5. The method of claim 4, wherein, The constructing the loss function of the image processing model according to the first distance and the second distance comprises: determining a predetermined distance interval; constructing the loss function of the image processing model according to the first distance, the second distance and the predetermined distance interval, wherein the loss function requires the first distance to be at least smaller than the second distance by the predetermined distance interval.

6. The method according to any one of claims 3 to 5, characterized in that, The third image is an image of a second local region in the target image.

7. The method of claim 6, wherein, The performing an alignment operation on the first image to obtain a first image with the same size as the target image to obtain a second image corresponding to the first image comprises: performing an alignment operation on the first image by using an alignment module in the image processing model to obtain a first image with the same size as the target image; extracting a first feature of the first image with the same size as the target image and extracting a second feature of the target image by using a feature extraction module in the image processing model; processing the first feature and the second feature by using a processing module in the image processing model to obtain a second image corresponding to the first image.

8. The method of claim 6, wherein, The obtaining a third image with a similarity to the second image lower than a predetermined threshold value comprises: obtaining a plurality of second local regions, wherein the intersection of the plurality of second local regions and the first local region is less than the predetermined threshold value; selecting a second local region from the plurality of second local regions, and taking an image of the selected second local region as the third image.

9. The method of claim 6, wherein, The method further comprises: obtaining an image of a local region in a fourth image and a fifth image; processing the image of the local region in the fourth image and the fifth image by using the optimized image processing model to obtain a corresponding region image in the fifth image corresponding to the local region in the fourth image.

10. The method of claim 9, wherein, The image processing model comprises an image recognition model for recognizing at least one of the following predetermined local regions: a local region in a pedestrian image comprising a person, a local region in an object image comprising an object, and a local region in a scene image comprising a person and an object.

11. The method of claim 10, wherein, The predetermined local region comprises a local region of an image frame intercepted from a monitoring image video.

12. A model processing method characterized by comprising: The method further comprises: receiving a model optimization instruction through an interactive interface; receiving a first image through the interactive interface based on the model optimization instruction; displaying a model optimization result on the interactive interface, wherein the model optimization result comprises an optimized image processing model, wherein the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, the second distance is a distance between a negative sample pair, the first image and a second image form the positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image by using the image processing model before optimization, the first image and a third image form the negative sample pair, the third image is an image with a similarity to the second image lower than a predetermined threshold value, and the image processing model is used to perform the model processing method in any one of claims 1 to 2.

13. A model processing method characterized by comprising: The method further comprises: receiving, by a front-end client, a first region image, wherein the first region image is an image of a local region of a first predetermined image; sending, by the front-end client, the first region image to a back-end server, and receiving a region processing result returned by the back-end server, wherein the region processing result is processed by an optimized image processing model, and the region processing result includes an image of a corresponding region of the first predetermined image in a second predetermined image, wherein the optimized image processing model is obtained by optimizing an image processing model according to a first distance and a second distance, the first distance is a distance between positive sample pairs, the second distance is a distance between negative sample pairs, a first image and a second image are a positive sample pair, the second image is an image corresponding to the first image obtained by processing the first image using the image processing model before optimization, the first image and a third image are a negative sample pair, the third image is an image with a similarity to the second image lower than a predetermined threshold, and the image processing model is used to perform the model processing method of any one of claims 1-2; displaying, by the front-end client, the region processing result.

14. A model processing apparatus characterized by comprising: comprising: a first obtaining module, configured to obtain a first pedestrian local region image in a first pedestrian image and a second pedestrian image, wherein the first pedestrian image is used to represent various images including a person, and the first pedestrian local region image is used to represent images of various parts of a pedestrian; a first processing module, configured to obtain a size of the second pedestrian image, perform an alignment operation on the first pedestrian local region image using an image processing model to obtain a first pedestrian local region image with the same size as the second pedestrian image, and obtain a second pedestrian local region image corresponding to the first pedestrian local region image in the second pedestrian image, wherein the image processing model is obtained by optimizing a first distance and a second distance, the first distance is a distance between positive sample pairs in image samples, the second distance is a distance between negative sample pairs in image samples, and the first distance is at least smaller than the second distance by a predetermined distance interval.

15. A model processing apparatus characterized by comprising: comprising: a second processing module, configured to process a first image using an image processing model to obtain a second image corresponding to the first image, wherein the first image and the second image are a positive sample pair; a second obtaining module, configured to obtain a third image with a similarity to the second image lower than a predetermined threshold, wherein the first image and the third image are a negative sample pair; a determining module, configured to determine a first distance between the positive sample pairs and a second distance between the negative sample pairs; an optimizing module, configured to optimize the image processing model according to the first distance and the second distance, wherein the image processing model is used to perform the model processing method of any one of claims 1-2.

16. A model processing apparatus characterized by comprising: comprising: a first receiving module, configured to receive a model optimization instruction through an interactive interface; a second receiving module, configured to receive a first image through the interactive interface based on the model optimization instruction. The first display module is configured to display a model optimization result in the interactive interface, wherein the model optimization result comprises an optimized image processing model, and the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, the second distance is a distance between a negative sample pair, the first image and the second image form the positive sample pair, the second image is an image corresponding to the first image and obtained by processing the first image using the image processing model before optimization, the first image and the third image form the negative sample pair, the third image is an image having a similarity lower than a predetermined threshold with the second image, and the image processing model is configured to perform the model processing method in any one of claims 1 to 2.

17. A model processing apparatus characterized by comprising: The method comprises: The third receiving module is configured to receive a first region image from a front-end client, wherein the first region image is an image of a local region of a first predetermined image; The fourth receiving module is configured to send the first region image from the front-end client to a back-end server and receive a region processing result returned by the back-end server, wherein the region processing result is obtained by processing the first region image using the optimized image processing model, and the region processing result comprises an image of a corresponding region of a second predetermined image, wherein the optimized image processing model is obtained by optimizing the image processing model according to a first distance and a second distance, the first distance is a distance between a positive sample pair, the second distance is a distance between a negative sample pair, the first image and the second image form the positive sample pair, the second image is an image corresponding to the first image and obtained by processing the first image using the image processing model before optimization, the first image and the third image form the negative sample pair, and the third image is an image having a similarity lower than a predetermined threshold with the second image; The second display module is configured to display the region processing result on the front-end client, and the image processing model is configured to perform the model processing method in any one of claims 1 to 2.

18. A storage medium, characterized by The storage medium comprises a stored program, wherein the program controls a device in which the storage medium is located to perform the model processing method in any one of claims 1 to 13 when the program is running.

19. A computer device, comprising: The method comprises: a memory and a processor, the memory stores a computer program; the processor is configured to execute the computer program stored in the memory, and the computer program is configured to enable the processor to perform the model processing method in any one of claims 1 to 13 when the computer program is running.

Citation Information

Patent Citations

  • A pedestrian re-identification method and device

    CN109784166A