The invention discloses an operation
interaction method and
system applied to
camera image editing, and relates to the technical field of
image processing, and the method comprises the steps: obtaining a voice semantic
heat map, a pointing intensity map, a touch
confidence map and a gazing
confidence map based on a multi-
modal interaction data packet, and calculating an image
feature matrix at the same time; fusing into a multi-
modal evidence graph through a normalized scale; performing semantic segmentation according to the image
feature matrix to obtain a semantic segmentation first draft and a pixel-by-pixel category confidence coefficient, and performing position correlation weighting on the pixel-by-pixel category confidence coefficient by taking the multi-
modal evidence graph as a confidence coefficient
modulation factor to generate a candidate object
mask sequence; and performing highlight display on the candidate object
mask sequence, and performing conflict resolution and priority rearrangement in combination with the multi-mode evidence graph to generate a target object
mask. According to the method, deep fusion of the interaction intention and
image segmentation is realized, the precision and consistency of candidate
region detection are improved, and the stability of real-time rendering and the reliability of an editing result are improved.