A method, device and electronic device for processing a perception task
By calculating the matching degree of perception tasks and image feature vectors and optimizing the processing of perception tasks, the problems of low resource utilization and low efficiency of traditional mobile collaborative perception systems are solved, and efficient perception task processing is achieved.
Patent Information
- Application Number
- CN202110005349.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-01-05
AI Technical Summary
When traditional mobile collaborative perception systems handle multiple perception tasks, the resource utilization rate is low, the overhead is high, and the efficiency is low, resulting in the perception nodes needing repeated work, and the system processing complexity is O(n).
By obtaining the feature vector and image feature vector corresponding to each perceptual task, compute the matching degree, and determine whether the image meets the needs of the perceptual task based on the matching degree, thereby optimizing resource utilization and system processing efficiency.
It realizes efficient processing of perceptual tasks, reduces the duplicate work of perceptual nodes, improves resource utilization and system processing efficiency, and reduces the system processing complexity to O(1).
Smart Images

Figure CN114780229B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to a method, device and electronic equipment for processing a perception task. Background Art
[0002] In the IoT system, perception is the entrance to data and the basis for the operation and service provision of the IoT system. In many fields such as industry, agriculture, medical care, traffic monitoring, climate monitoring, smart cities, etc., perception is the basic function of the IoT system. In the past, when deploying perception systems in each field, sensors in that field were deployed separately. When security and traffic issues needed to be controlled in the city, a large number of cameras were deployed; when manhole covers needed to be monitored, sensors were deployed under the manhole covers; when environmental pollution needed to be monitored, environmental monitoring sensors were deployed in the city; when urban noise needed to be detected, noise monitoring sensors were deployed in the city, and so on. It is extremely costly to deploy a completely independent perception system and new perception equipment. Moreover, when there are more and more perception systems in the city, maintenance becomes difficult and the cost of later maintenance is very high. In fact, the sensors and public mobile devices already deployed in the city, such as the mobile phones of city patrol personnel and staff, can collaboratively complete a large number of perception tasks in the city, that is, mobile collaborative perception.
[0003] As for traditional mobile collaborative sensing, from the traditional method, the mobile collaborative sensing system needs to collect different data sets for different sensing tasks, that is, n sensing tasks correspond to n data sets. For n sensing tasks, if a sensing node is selected as the input of m sensing tasks, the node needs to work m times, the work of the sensing node is repeatedly consumed, and the resource usage and system processing efficiency are low. Although compared with the redeployment of sensors, the efficiency of mobile collaborative sensing has been greatly improved and the cost has also decreased, but from a system perspective, this model is still expensive and inefficient. Summary of the invention
[0004] The present invention provides a method, device and electronic equipment for processing a perception task, so as to improve the processing efficiency of the perception task.
[0005] To solve the above technical problems, the embodiments of the present invention provide the following solutions:
[0006] A method for processing a perception task, comprising:
[0007] Obtaining a first eigenvector corresponding to each perception task and a second eigenvector corresponding to the first image;
[0008] Calculating a matching degree between the second feature vector and the first feature vector corresponding to each perception task, wherein a target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task;
[0009] According to the matching degree, it is determined whether the first image meets the task requirements of the perception task, wherein, if the first image meets the task requirements, the first image includes the target object required by the perception task.
[0010] Optionally, determining whether the first image meets the task requirement of the perception task according to the matching degree includes:
[0011] According to the matching degree, assigning a first weight corresponding to the perception task to the first image through an attention mechanism, wherein the first weight assigned by the attention mechanism is positively correlated with the target feature;
[0012] When the first weight is greater than or equal to the requirement weight corresponding to the perception task, it is confirmed that the first image meets the task requirement of the perception task.
[0013] Optionally, after assigning a first weight corresponding to the perception task to the first image according to the matching degree, the method further includes:
[0014] When the first weight is less than the requirement weight corresponding to the perception task, it is confirmed that the first image does not meet the task requirement of the perception task.
[0015] Optionally, after confirming that the first image meets the task requirements of the perception task, the method further includes:
[0016] generating a copy image of the first image;
[0017] After adding a first tag to the image data of the first image, the first image is placed into a resource pool, wherein the first tag is used to indicate that the first image meets the task requirements of the perception task.
[0018] Optionally, after confirming that the first image does not meet the task requirement of the perception task, the method further includes:
[0019] The first image is stored in the resource pool.
[0020] Optionally, the method further includes:
[0021] Obtaining the timestamp of the second image in the resource pool;
[0022] When the time indicated by the timestamp is not within the task requirement time range of any current perception task, the second image is discarded from the resource pool.
[0023] Optionally, after obtaining the timestamp of the second image in the resource pool, the method further includes:
[0024] When the time indicated by the timestamp falls within the task requirement time range of N perception tasks and the proportion of the second image determined to meet the task requirements in the N perception tasks is less than a threshold, the second image is discarded from the resource pool, where N represents a positive integer.
[0025] Optionally, the method further includes:
[0026] If any perception task fails to confirm that the second image meets the task requirement, the second image is discarded from the resource pool.
[0027] Optionally, obtaining a second feature vector corresponding to the first image includes:
[0028] Normalizing each pixel in the first image to obtain image information feature points;
[0029] The image information feature points are extracted through the feature extraction model, and the image information feature points are concatenated into a second feature vector.
[0030] The embodiment of the present invention further provides a processing device for a perception task, comprising:
[0031] A first acquisition module, used to acquire a first feature vector corresponding to each perception task and a second feature vector corresponding to the first image;
[0032] a calculation module, used for calculating the matching degree between the second feature vector and the first feature vector corresponding to each perception task, wherein the target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task;
[0033] The confirmation module is used to determine whether the first image meets the task requirements of the perception task according to the matching degree, wherein, if it meets the task requirements, the first image includes the target object required by the perception task.
[0034] The present invention also provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the method as described above through the computer program.
[0035] The present invention also provides a processor-readable storage medium, which stores processor-executable instructions. The processor-executable instructions are used to enable the processor to execute the method as described above.
[0036] The above solution of the present invention includes at least the following beneficial effects:
[0037] The above scheme of the present invention obtains the first feature vector corresponding to each perception task and the second feature vector corresponding to the first image; calculates the matching degree between the second feature vector and the first feature vector corresponding to each perception task, wherein the target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task; and determines whether the first image meets the task requirements of the perception task based on the matching degree. The method of the embodiment of the present invention can perform feature extraction on the image data, and calculate the matching degree between the feature vector corresponding to the perception task and the feature vector corresponding to the image, wherein the matching degree is positively correlated with the target feature, so that the important features in the image that match the task requirements can be highlighted according to the requirements of different perception tasks, so as to accurately calculate the matching degree between the feature vector corresponding to the image and the feature vector corresponding to each perception task, thereby improving resource utilization and system processing efficiency, reducing duplication of perception nodes, and improving the processing efficiency of perception tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A flowchart of a method for processing a perception task according to an embodiment of the present invention;
[0039] Figure 2 A flowchart of a method for processing a perception task according to another embodiment of the present invention;
[0040] Figure 3 A flowchart of a method for processing a perception task according to another embodiment of the present invention;
[0041] Figure 4 A schematic diagram of feature extraction according to an embodiment of the present invention;
[0042] Figure 5 is a schematic diagram of a first image according to an embodiment of the present invention;
[0043] Figure 6 is a schematic diagram of processing a first image according to an embodiment of the present invention;
[0044] Figure 7 A schematic diagram of a module of a device for processing a perception task according to an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0046] In a mobile collaborative perception system, the system can collaborate with fixed-deployed perception devices and mobile perception nodes (such as robots) to collect perception information. The amount and type of data are increasing, so it is necessary to make full use of perception resources and perception information.
[0047] The image information collected by the existing perception systems contains {Ai, Bi, Ci, Di…} types. For example, by analyzing the pictures or videos collected by the city cameras, we can analyze the road cracks, whether the manhole cover is missing, traffic accidents, pedestrian flow and other information. At the same time, the camera also contains the location information of its equipment installation and the time information of data collection. For the traditional mobile collaborative perception method, each task within the same task time needs to match the corresponding perception node. For a perception node, if it needs to serve n tasks, the node needs to work n times, and the complexity of the system processing is O(n).
[0048] like Figure 1 As shown, an embodiment of the present invention provides a method for processing a perception task, including:
[0049] Step 11, obtaining a first eigenvector corresponding to each perception task and a second eigenvector corresponding to the first image;
[0050] Step 12, calculating the matching degree between the second feature vector and the first feature vector corresponding to each perception task, wherein the target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task;
[0051] In an embodiment of the present invention, there is a positive correlation between the target feature and the matching degree, so that the target feature in the first image that matches the perception task can be highlighted, so that the calculation of the matching degree focuses more on the target feature in the first image, reducing the influence of other features on the matching degree calculation, and improving the accuracy of the matching. In an optional embodiment of the present invention, the matching degree can be calculated by an attention mechanism.
[0052] Step 13: Determine whether the first image meets the task requirements of the perception task based on the matching degree, wherein if the first image meets the task requirements, the first image includes the target object required for the perception task.
[0053] The method of the embodiment of the present invention can extract features from image data and calculate the matching degree between the feature vector corresponding to the perception task and the feature vector corresponding to the image. The matching degree here is positively correlated with the target feature, so that important features in the image that match the task requirements can be highlighted according to the requirements of different perception tasks, thereby accurately calculating the matching degree between the feature vector corresponding to the image and the feature vector corresponding to each perception task, thereby improving resource utilization and system processing efficiency, reducing duplication of work of perception nodes, and improving the processing efficiency of perception tasks.
[0054] In the method of the embodiment of the present invention, the matching degree between the second eigenvector corresponding to the first image and the first eigenvector corresponding to each perception task is calculated, and whether the first image meets the task requirements of the perception task is determined based on the matching degree, so that the processing of the first image can be adapted to each perception task. Through the method of the embodiment of the present invention, the perception node only needs to work once to serve n tasks within the task time, and the complexity of system processing is O(1). The perception node can reduce a large amount of workload, the system can reduce the computing load, and the work efficiency of the system is greatly optimized. It solves the problems of low resource utilization, high overhead, and low efficiency in the processing of mobile collaborative perception tasks.
[0055] In an optional embodiment of the present invention, determining whether the first image meets the task requirements of the perception task based on the degree of matching includes: assigning a first weight corresponding to the perception task to the first image through an attention mechanism based on the degree of matching, wherein the first weight assigned by the attention mechanism is positively correlated with the target feature; and confirming that the first image meets the task requirements of the perception task when the first weight is greater than or equal to the required weight corresponding to the perception task. In an embodiment of the present invention, when assigning the first weight corresponding to the perception task to the first image, the weight of the elements in the first image that match the perception task can be highlighted through the attention mechanism, so that the more elements in the first image that match the perception task, the larger the first weight assigned, so that the assignment of the first weight is positively correlated with the features in the first image that match the perception task. The method of the embodiment of the present invention can assign the first weight to the first image according to different perception tasks by using the attention mechanism, thereby improving the accuracy of judging whether the first image meets the task requirements of the perception task.
[0056] In an optional embodiment of the present invention, after assigning a first weight corresponding to the perception task to the first image based on the degree of matching, the method further includes: when the first weight is less than the required weight corresponding to the perception task, confirming that the first image does not meet the task requirements of the perception task.
[0057] In an embodiment of the present invention, a first weight corresponding to the perception task is assigned to the first image, so that the first weight is compared with the requirement weight corresponding to the perception task. When the first weight is greater than or equal to the requirement weight, it is confirmed that the first image meets the requirement of the perception task; when the first weight is less than the requirement weight, it is confirmed that the first image does not meet the requirement of the perception task. This allows a more accurate grasp of whether the first image meets the task requirement of each perception task. The method can be applicable to different perception tasks, and the value of the method of the embodiment of the present invention is more obvious in particular for complex perception systems.
[0058] In an optional embodiment of the present invention, after confirming that the first image meets the task requirements of the perception task, the method also includes: generating a copy image of the first image; after adding a first mark to the image data of the first image, placing the first image into a resource pool, wherein the first mark is used to indicate that the first image meets the task requirements of the perception task.
[0059] In an optional implementation manner of the present invention, after confirming that the first image does not meet the task requirements of the perception task, the method further includes: storing the first image in a resource pool.
[0060] In the embodiment of the present invention, a resource pool is introduced to solve the problem of excessive storage and computing overhead that may be caused by repeated introduction of data.
[0061] In an optional embodiment of the present invention, the method further includes: obtaining a timestamp of the second image in the resource pool; and discarding the second image from the resource pool if the time indicated by the timestamp is not within the task requirement time range of any current perception task.
[0062] In an optional embodiment of the present invention, after obtaining the timestamp of the second image in the resource pool, the method further includes: discarding the second image from the resource pool when the time indicated by the timestamp belongs to the task requirement time range of N perception tasks and the proportion of the second images determined in the N perception tasks that meet the task requirements is less than a threshold, wherein N represents a positive integer.
[0063] In an optional implementation manner of the present invention, the method further includes: if any perception task fails to confirm that the second image meets the task requirements, discarding the second image from the resource pool.
[0064] In the embodiment of the present invention, if the data is kept in the data pool, the amount of data in the data pool will become larger and larger, and the computing resources and storage resources of the system will be consumed very much. In order to avoid this problem, it is determined whether to discard the image by judging whether to discard it, so as to avoid the worthless data being kept in the resource pool.
[0065] In an optional embodiment of the present invention, obtaining a second feature vector corresponding to the first image includes: normalizing each pixel in the first image to obtain image information feature points; extracting the image information feature points through a feature extraction model, and splicing the image information feature points into a second feature vector.
[0066] Combine the following Figures 2 to 6 , the method of the embodiment of the present invention is illustrated by way of example.
[0067] like Figures 2 to 3 As shown, the method in the embodiment of the present invention includes:
[0068] Step 1: For the mobile collaborative perception system, according to the system requirements, first extract the demand feature vector vector.demand for the task requirements of the perception task. For example, if Task 1 needs to collect road cracks in the city, then the task requirements need to be input in advance and the demand vector vector.crack needs to be generated.
[0069] Step 2: After the image is collected, the image is processed using an image model, such as the VGG model (Visual Geometry Group Network), to extract the high-dimensional features vector.image of the image data, and the system adds the collection timestamp ti to the image information. After extracting the high-dimensional feature information of the image, the attention mechanism calculates the matching degree between the high-dimensional feature vector vector.image of the image and the task demand vector vector.demand, and assigns a weight corresponding to the task to vector.image.
[0070] Step 3: According to the requirements of different tasks for collected images, when the weight hi calculated by the attention mechanism is less than the required weight Hi, the image data of the task does not meet the task requirements, the data directly enters the resource pool, and the system marks the data with a timestamp; when the weight hi calculated by the attention mechanism is greater than or equal to the required weight Hi, the image data is considered to meet the task requirements, the system immediately generates a copy of the data and keeps it under the corresponding task, and adds the identifier task to the original data i ; and put the original data into the resource pool. That is, after image acquisition and weights are assigned to the acquired images for different tasks, all data will enter the resource pool, and only the data that is valuable to each task will be retained.
[0071] Step 4: If the data is kept in the data pool, the amount of data in the data pool will increase, and the system's computing resources and storage resources will be consumed very much. To avoid this problem, a data value judgment mechanism is added to the resource pool. Whenever data enters the resource pool, the resource pool immediately judges the value of the data. The judgment method is as follows:
[0072] 1) Determine based on timestamp ti:
[0073] a) Check whether the timestamp ti of the data belongs to the time range of any task in the mobile collaborative sensing system. If not, the data is discarded;
[0074] b) If the timestamp belongs to the time range of some tasks, but has been judged as worthless data in X% of these tasks, the data is discarded; here X can be a natural number greater than 50, such as 70-85.
[0075] 2) Determine based on data utilization:
[0076] a) After a cycle, the data is not used by any task, that is, there is no task in the data i label; then the data is judged as worthless data and the system discards the data;
[0077] After the resource pool determines the value of the data, it divides the data into two categories: valuable data and worthless data.
[0078] Step 5: Feed back valuable data into the image information and provide information for the continuously updated perception task together with the newly collected images. The worthless data is directly discarded.
[0079] See also Figure 4 In the embodiment of the present invention, for the task requirements and the collected image information processing, each pixel can be normalized first: image = u / 225; then VGG is used to extract the normalized image information feature points, and all the feature points are spliced into a high-dimensional feature vector.
[0080] After the above steps, we can obtain the high-dimensional feature vector vector.demand of the task requirement and the high-dimensional feature vector.image of the collected image information.
[0081] In the method of the embodiment of the present invention, the extracted high-dimensional feature vector is processed by the attention mechanism, and the feature values in each type of information are weighted according to the requirements of different perception tasks, so that the feature value image of each type of information and the task requirement D can be obtained. i The relationship between them is as follows:
[0082] task i Task requirements D i = <key i ,d i >, where key i Yes i The weight coefficient corresponds to the demand factor of the i-th task, d i represents the demand of the i-th task;
[0083] Image Data:
[0084] Among them, vector.image i Represents the high-dimensional feature vector of image i, vector.d i represents the feature vector of task i, and n is the number of each task. The calculation process of the attention mechanism is as follows:
[0085] 1) Use the multi-layer neural network model MLP (Multi-Layer Perceptron) to calculate the correlation between di and each high-dimensional feature vector:
[0086] similarity(vectori.image,d i ) = MLP(vector.image,d i )
[0087] 2) Normalize the scores calculated in 1) through the softmax of the logistic regression model to highlight the weights of important elements:
[0088]
[0089] where hi represents the weight calculated by the attention mechanism, and sim i represents the value of the similarity function in 1);
[0090] 3) After calculating the weights for each requirement of each task through the above two steps, weighted summation can obtain the attention values matching each requirement of different tasks for different data:
[0091]
[0092] where hi represents the value of the softmax function after normalization; d i is the requirement of the i-th task. When setting the initial weight Ii of the system as the discrimination basis:
[0093] When attention(vector.image,vector.di) < Ii, the data will be returned to the data pool as a common resource that can provide information for other tasks; when attention(vector,D i ) ≥ I i ), a copy of the data is retained in the task resource pool, and the original data is returned to the data pool as a common resource that can provide information for other tasks.
[0094] When all the data returns to the resource pool, to avoid continuous increase of the system's storage and computing resources and overloading of the system, the resource pool will perform a value determination on the data:
[0095] a) Whether the timestamp ti carried by the data belongs to the time range of any task in the mobile collaborative perception system. If not, the data is discarded; taski_t is the task time corresponding to taski;
[0096] b) If the timestamp belongs to the time range of some tasks, but there is a history of being determined as worthless data in Y% of these tasks, the data is discarded; here Y can take natural numbers greater than 50, such as 70 - 85.
[0097] c) After one cycle, the data is not used by any task, that is, there is no task in the data i label; then the data is judged as worthless data and the system discards the data;
[0098] The method of the embodiment of the present invention saves system overhead in the stage of perceptual information collection. It does not need to create a new perceptual task for each perceptual task and let the corresponding perceptual node perform the task. Instead, the perceptual information collected by the existing perceptual nodes within the same task time is reused multiple times. At the same time, the embodiment of the present invention improves the attention mechanism, introduces a resource pool and a data value judgment process, and solves the problem of excessive storage and computing overhead caused by repeated data introduction.
[0099] See also Figure 5-6 , the present invention is illustrated by taking road problem detection as an example.
[0100] For the information collected in the perception system, Figure 5 As shown in the picture information, when the mobile collaborative perception system includes the tasks of collecting manhole cover information, road damage / crack information, and road obstacle information, the mobile collaborative perception system needs to issue three different tasks, and the perception node needs to perform three tasks. Using the method of the embodiment of the present invention, the perception node only needs to collect data once, and the system can calculate the information required for the corresponding task from the collected data according to the different requirements of the three tasks. Figure 6 As shown, Figure 5 When the pictures shown are matched with different tasks, it can be confirmed whether they meet the task requirements according to different tasks, for example Figure 6 The lower left image in the picture shown can be matched to the task of manhole cover recognition; Figure 6 The lower middle picture in the image shown can be matched to the task of road surface depression / crack identification; Figure 6 The picture in the lower right corner of the picture shown can match the task of road obstacle recognition. The method of the embodiment of the present invention improves the processing mode of one task distribution and data collection corresponding to one perception task in traditional mobile collaborative perception by adding a resource pool and a data copy in the attention mechanism, and increases the available data resources for the traditional target recognition method.
[0101] In addition, the method of the embodiment of the present invention can also be applied to the security monitoring of MMS content. Many people use the collection and release of advertisements containing bad information, which has a negative impact on the lives of users. For this, it is necessary to establish multiple perception tasks to perceive whether the MMS involves bad information. Through the method of the embodiment of the present invention, the first feature vector corresponding to each perception task and the second feature vector corresponding to the MMS information can be obtained, and the matching degree between the second feature vector and the first feature vector corresponding to each perception task is calculated through the attention mechanism; according to the matching degree, it is determined whether the MMS information meets the task requirements of the perception task, so that the MMS information transmitted by the user can be effectively analyzed in multiple dimensions and divided into different categories, reducing the impact of bad information on user experience and the reputation of the operator service provider.
[0102] The method of the embodiment of the present invention can also be applied to content security review. Digital content providers and operating entities manage and operate massive amounts of digital information and provide short video services to users. For the emerging short video market, the platform's image review strategy is still not sound, and many bloggers upload bad information that violates platform rules in the form of playing edge balls. For operators, the control of such information should be more strict and cautious, and multiple perception tasks need to be established, each perception task is used to perceive whether digital information (such as video data, image data) complies with different platform rules. Through the method of the embodiment of the present invention, the first eigenvector corresponding to each perception task and the second eigenvector corresponding to the digital information can be obtained, and the matching degree between the second eigenvector and the first eigenvector corresponding to each perception task is calculated by the attention mechanism; according to the matching degree, it is determined whether the digital information meets the task requirements of the perception task, so that video data, picture information, etc. can be cyclically analyzed from multiple dimensions, and the processing efficiency and accuracy of the platform can be greatly improved, which can eliminate adverse social impacts and improve service quality and customer trust.
[0103] The method of the embodiment of the present invention can also be applied to camera information analysis. As an important tentacle of current security, cameras have been widely used throughout the country. The pictures collected by the camera contain a large amount of information, such as urban security, road safety, pedestrian and vehicle flows, etc. Different perception tasks need to be established to perceive urban security, road safety, pedestrian and vehicle flows, etc. respectively. Through the method of the embodiment of the present invention, the first eigenvector corresponding to each perception task and the second eigenvector corresponding to the image collected by the camera can be obtained, and the matching degree between the second eigenvector and the first eigenvector corresponding to each perception task is calculated through the attention mechanism; according to the matching degree, it is determined whether the picture meets the task requirements of the perception task, so that various problems in the city or community can be collaboratively processed, which improves the efficiency of security management and information collection / processing.
[0104] like Figure 7As shown, the embodiment of the present invention further provides a processing device 70 for a perception task, comprising:
[0105] A first acquisition module 71, used to acquire a first feature vector corresponding to each perception task and a second feature vector corresponding to the first image;
[0106] A calculation module 72 is used to calculate the matching degree between the second feature vector and the first feature vector corresponding to each perception task, wherein the target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task;
[0107] The confirmation module 73 is used to determine whether the first image meets the task requirements of the perception task according to the matching degree, wherein, if it meets the task requirements, the first image includes the target object required by the perception task.
[0108] Optionally, the confirmation module 73 may include:
[0109] an allocating unit, configured to allocate a first weight corresponding to the perception task to the first image through an attention mechanism according to the matching degree, wherein the first weight allocated by the attention mechanism is positively correlated with the target feature;
[0110] The first confirmation unit is used to confirm that the first image meets the task requirement of the perception task when the first weight is greater than or equal to the requirement weight corresponding to the perception task.
[0111] Optionally, the confirmation module 73 may further include:
[0112] The second confirmation unit is used to confirm that the first image does not meet the task requirement of the perception task when the first weight is less than the requirement weight corresponding to the perception task.
[0113] Optionally, the device 70 may further include:
[0114] A generating module, configured to generate a copy image of the first image;
[0115] The processing module is used to add a first mark to the image data of the first image and then put the first image into a resource pool, wherein the first mark is used to indicate that the first image meets the task requirements of the perception task.
[0116] Optionally, the device 70 may further include:
[0117] The storage module is used to store the first image into the resource pool.
[0118] Optionally, the device 70 may further include:
[0119] A second acquisition module, used for acquiring a timestamp of a second image in the resource pool;
[0120] The first discarding module is used to discard the second image from the resource pool when the time indicated by the timestamp is not within the task requirement time range of any current perception task.
[0121] Optionally, the device 70 may further include:
[0122] The second discarding module is used to discard the second image from the resource pool when the time indicated by the timestamp belongs to the task requirement time range of N perception tasks and the proportion of the second images determined in the N perception tasks that meet the task requirements is less than a threshold, wherein N represents a positive integer.
[0123] Optionally, the device 70 may further include:
[0124] The third discarding module is used to discard the second image from the resource pool when any perception task fails to confirm that the second image meets the task requirements.
[0125] Optionally, the first acquisition module 71 may include:
[0126] A normalization unit, used to perform normalization processing on each pixel point in the first image to obtain image information feature points;
[0127] The splicing unit is used to extract image information feature points through a feature extraction model, and splice the image information feature points into a second feature vector.
[0128] It should be noted that the device is a device corresponding to the above method embodiment, and all implementation methods in the above method embodiment are applicable to the embodiment of the device and can achieve the same technical effect.
[0129] The embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above method through the computer program. All implementations in the above method embodiment are applicable to the embodiment of the device, and can also achieve the same technical effect.
[0130] The embodiment of the present invention further provides a processor-readable storage medium, wherein the processor-readable storage medium stores processor-executable instructions, wherein the processor-executable instructions are used to cause the processor to execute the method described above. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0131] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0133] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0134] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0136] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical disks.
[0137] In addition, it should be noted that in the apparatus and method of the present invention, it is obvious that each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. Moreover, the steps of performing the above-mentioned series of processing can naturally be performed in chronological order according to the order of description, but it is not necessary to perform them in chronological order, and some steps can be performed in parallel or independently of each other. For those of ordinary skill in the art, it is understood that all or any steps or components of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or a network of computing devices in hardware, firmware, software or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.
[0138] Therefore, the purpose of the present invention can also be achieved by running a program or a group of programs on any computing device. The computing device can be a well-known general device. Therefore, the purpose of the present invention can also be achieved by simply providing a program product containing a program code that implements the method or device. That is to say, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be pointed out that in the device and method of the present invention, it is obvious that each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. In addition, the steps of performing the above-mentioned series of processing can naturally be performed in chronological order according to the order of description, but it is not necessary to perform them in chronological order. Some steps can be performed in parallel or independently of each other.
[0139] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for processing a perception task, characterized in that: include: Obtaining a first eigenvector corresponding to each perception task and a second eigenvector corresponding to the first image; Calculating a matching degree between the second feature vector and the first feature vector corresponding to each of the perception tasks, wherein a target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task; According to the matching degree, it is determined whether the first image meets the task requirements of the perception task, wherein, if the first image meets the task requirements, the first image includes the target object required by the perception task.
2. The method according to claim 1, characterized in that Determining whether the first image meets the task requirement of the perception task according to the matching degree includes: According to the matching degree, assigning a first weight corresponding to the perception task to the first image through an attention mechanism, wherein the first weight assigned by the attention mechanism is positively correlated with the target feature; When the first weight is greater than or equal to the requirement weight corresponding to the perception task, it is confirmed that the first image meets the task requirement of the perception task.
3. The method according to claim 2, characterized in that After assigning a first weight corresponding to the perception task to the first image according to the matching degree, the method further includes: When the first weight is less than the requirement weight corresponding to the perception task, it is confirmed that the first image does not meet the task requirement of the perception task.
4. The method according to claim 2, characterized in that: After confirming that the first image meets the task requirements of the perception task, the method further includes: generating a copy image of the first image; After adding a first mark to the image data of the first image, the first image is placed in a resource pool, wherein the first mark is used to indicate that the first image meets the task requirement of the perception task.
5. The method according to claim 3, characterized in that: After confirming that the first image does not meet the task requirements of the perception task, the method further includes: The first image is stored in a resource pool.
6. The method according to claim 4 or 5, characterized in that: Also includes: Acquire a timestamp of a second image in the resource pool; When the time indicated by the timestamp is not within the task requirement time range of any current perception task, the second image is discarded from the resource pool.
7. The method according to claim 6, characterized in that After acquiring the timestamp of the second image in the resource pool, the method further includes: When the time indicated by the timestamp falls within the task requirement time range of N perception tasks and the proportion of the second image determined to meet the task requirements in the N perception tasks is less than a threshold, the second image is discarded from the resource pool, where N represents a positive integer.
8. The method according to claim 4 or 5, characterized in that: Also includes: If any of the perception tasks fails to confirm that the second image meets the task requirements, the second image is discarded from the resource pool.
9. The method according to claim 1, characterized in that: Obtaining a second feature vector corresponding to the first image includes: Performing normalization processing on each pixel in the first image to obtain image information feature points; The image information feature points are extracted through a feature extraction model, and the image information feature points are concatenated into the second feature vector.
10. A processing device for a perception task, characterized in that: include: A first acquisition module, used to acquire a first feature vector corresponding to each perception task and a second feature vector corresponding to the first image; a calculation module, configured to calculate a matching degree between the second feature vector and a first feature vector corresponding to each of the perception tasks, wherein a target feature is positively correlated with the matching degree, and the target feature is a feature in the first image that matches the perception task; A confirmation module is used to determine whether the first image meets the task requirements of the perception task based on the matching degree, wherein, if the first image meets the task requirements, the first image includes the target object required by the perception task.
11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 9 through the computer program.
12. A processor-readable storage medium, characterized in that: The processor-readable storage medium stores processor-executable instructions, and the processor-executable instructions are used to enable the processor to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Data processing method and device
CN105488044A
Multimedia crowd sensing excitation method for a machine learning system
CN109872058A