Mobile terminal edge intelligent reasoning method and system supporting multi-modal input

By generating target recognition areas and combining them with multimodal data analysis strategies, the problem of difficulty in parsing multimodal data in existing technologies has been solved, enabling accurate identification of participants' needs and service adaptation, thus improving the meeting experience.

CN120997879AActive Publication Date: 2025-11-21NORTHERN INST OF AUTOMATIC CONTROL TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511104808.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-21
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing technologies struggle to comprehensively analyze the multimodal data of participants during meetings, making it impossible to accurately identify their needs and affecting the adaptability and timeliness of services.

Method used

By identifying the static planning area and dynamic stitching area in the target image, a target recognition area is generated. Combined with multimodal data analysis strategies, the active adjustment information and triggered adjustment information of the participants are obtained, so as to achieve comprehensive and accurate processing of multimodal data.

Benefits of technology

It improves the accuracy and comprehensiveness of identifying the needs of participants, enabling precise provision of tailored services and enhancing the meeting experience for participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997879A_ABST
    Figure CN120997879A_ABST
Patent Text Reader

Abstract

The invention provides a mobile terminal edge intelligent reasoning method and system supporting multi-modal input, and relates to a data processing technology, and the method comprises the steps: receiving a target image collected by a monitoring device for a target region through responding to front completion information, recognizing a static planning region and a dynamic splicing region of each target element in the target image, and splicing the regions, obtaining a target identification area of the target element, taking the corresponding target identification area as a trigger identification area based on the target amplitude of the target element in the target identification area, calling multi-modal data of the trigger identification area, and performing data analysis on the input data to obtain active adjustment information corresponding to the input data; the dynamic changes of the main body elements and the accompanying elements in the output data are identified based on the change identification strategy, the trigger adjustment information of the output data is obtained, the active adjustment information and the trigger adjustment information are sent to the adjustment equipment of the target elements, and multi-modal data can be analyzed and processed. And the demand information of the participants can be identified more comprehensively and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a mobile terminal edge intelligent inference method and system supporting multi-modal input. BACKGROUND

[0002] With the wide popularity of mobile devices and the rapid development of artificial intelligence technology, mobile terminal edge intelligent inference plays an increasingly key role in many fields. For example, in a conference scenario, through edge-side intelligent inference, various needs of conference participants can be accurately identified, conference quality is improved, and technical support is provided for the intelligent development of conferences.

[0003] However, the existing technology for identifying and processing personnel demand data has obvious defects in intelligent inference of data. For example, most existing methods analyze and infer single-modal data, and often cannot comprehensively analyze multi-modal data of personnel. For example, in an actual conference scenario, when facing various types of data such as voice data, video data, and text data of conference participants, the existing technology may be difficult to analyze and process these different modal data, may not be able to fully understand the demand signals conveyed by conference participants in different ways, and thus may not be able to accurately identify various needs of conference participants during the conference, resulting in inability to timely process the needs of conference participants.

[0004] Therefore, how to realize the analysis and processing of multi-modal data and more comprehensively and accurately identify the demand information of conference participants has become a problem to be solved. SUMMARY

[0005] The present application provides a mobile terminal edge intelligent inference method and system supporting multi-modal input, which can realize the analysis and processing of multi-modal data and more comprehensively and accurately identify the demand information of conference participants.

[0006] In a first aspect, the present application provides a mobile terminal edge intelligent inference method supporting multi-modal input, comprising: In response to the preposition completion information, receiving a target image collected by a monitoring device in a target area, identifying a static planning area and a dynamic splicing area of each target element in the target image and splicing to obtain a target recognition area of the target element; Based on the target amplitude of the target element in the target recognition area, the corresponding target recognition area is taken as a trigger recognition area, and multi-modal data of the trigger recognition area is called, the multi-modal data including input data and output data; Performing data analysis on the input data to obtain active adjustment information corresponding to the input data, identifying the dynamic changes of the subject element and the accompanying element in the output data based on a change identification strategy, and obtaining trigger adjustment information of the output data; a regulating device of the target element.

[0007] Optionally, in a possible implementation manner of the first aspect, the identifying the static planning area and the dynamic splicing area of each target element in the target image and splicing to obtain the target identification area of the target element comprises: identifying a position contour of a static position element and a placement contour of a static placement element in the target image; determining a processing contour of each placement side length as a position contour closest to the placement contour; processing the fixed side length and the extended side length in the processing contour to obtain the static planning area of the target element; mirroring and copying the static planning area based on the placement side length to obtain a copied planning area; performing dynamic supplement processing on the copied planning area according to an accompanying contour of an accompanying placement element in the placement contour to obtain the dynamic splicing area of the target element; splicing the static planning area and the dynamic splicing area to obtain the target identification area of the target element.

[0008] Optionally, in a possible implementation manner of the first aspect, the processing the fixed side length and the extended side length in the processing contour to obtain the static planning area of the target element comprises: regarding a side length parallel to a corresponding placement side length in the processing contour as a fixed side length, and regarding the remaining side lengths as extended side lengths; deleting the fixed side length closest to the corresponding placement side length, and extending the extended side length to one side of the corresponding placement side length until the extended side length intersects with the placement side length, to obtain the static planning area of the target element.

[0009] Optionally, in a possible implementation manner of the first aspect, the performing dynamic supplement processing on the copied planning area according to an accompanying contour of an accompanying placement element in the placement contour to obtain the dynamic splicing area of the target element comprises: identifying an accompanying contour of an accompanying placement element in the placement contour, regarding an intersection contour of the accompanying contour and the copied planning area as a merged contour, and regarding the remaining accompanying contours as candidate contours; obtaining a first combined contour according to a union set of the merged contour and the corresponding copied planning area; performing coordinate processing on the target image to obtain a straight line distance between a center point of the candidate contour and the first combined contour, and selecting a candidate contour with a straight line distance less than a preset screening distance as a selected contour of the corresponding first combined contour; Obtaining coordinate extreme values of the selected contour and the first combined contour, constructing a reconstruction rectangle based on the coordinate extreme values, and taking the reconstruction rectangle as a dynamic splicing area of the target element.

[0010] Optionally, in a possible implementation manner of the first aspect, the taking, as the trigger identification area, the corresponding target identification area based on the target amplitude of the target element in the target identification area comprises:

[0011] Optionally, in a possible implementation manner of the first aspect, the identifying the dynamic change of the subject element and the accompanying element in the output data based on the change identification strategy to obtain the trigger adjustment information of the output data comprises: obtaining output data, the output data comprising video data and audio data; identifying the dynamic change of the subject element and the accompanying element in the audio data based on an audio identification strategy to obtain trigger adjustment information of the audio data; identifying the dynamic change of the subject element and the accompanying element in the video data based on a video identification strategy to obtain trigger adjustment information of the video data.

[0012] Optionally, in a possible implementation manner of the first aspect, the identifying the dynamic change of the subject element and the accompanying element in the audio data based on the audio identification strategy to obtain the trigger adjustment information of the audio data comprises: identifying a subject sound wave in the audio data as the subject element and the rest sound waves as the accompanying elements based on tone; extracting a subject amplitude of the subject element in the audio data, and determining, when the subject amplitude is greater than a preset subject amplitude, a corresponding sound wave as a subject change sound wave and obtaining a subject duration of each subject change sound wave; determining, when the subject duration is less than a preset duration, the corresponding subject change sound wave as a transient subject sound wave, and retrieving preset transmission information based on the transient subject sound wave; determining, when the subject duration is greater than the preset duration, the corresponding subject change sound wave as a continuous subject sound wave, extracting text information corresponding to the continuous subject sound wave, and generating labeling adjustment information based on the text information; extracting an accompanying amplitude of the accompanying element in the audio data, and determining, when the accompanying amplitude is greater than a preset accompanying amplitude, a corresponding sound wave as an accompanying change sound wave and obtaining an accompanying duration of each accompanying change sound wave; determining, when the accompanying duration is less than a preset duration, the corresponding accompanying change sound wave as a transient accompanying sound wave, and retrieving disposal adjustment information based on the transient accompanying sound wave; ​When the accompanying duration is greater than the preset duration, the accompanying change sound wave is determined as a continuous accompanying sound wave, an elimination sound wave opposite to the continuous accompanying sound wave is generated, and the elimination sound wave is used as elimination adjustment information.

[0013] Optionally, in a possible implementation manner of the first aspect, the dynamic change of the subject element and the accompanying element in the video data according to the video recognition strategy comprises: identifying a target element in the video data as the subject element and an accompanying placed element as the accompanying element; identifying a change state of an element part corresponding to the subject element in the video data, and generating instruction adjustment information based on the change state; obtaining a display attribute of the accompanying element in the video data, the display attribute comprising a transparent attribute and a hidden attribute; when the display attribute is the transparent attribute, the corresponding accompanying element is determined as a recognizable element, a use height of the recognizable element is identified, and when the use height is less than a limit height, replacement adjustment information is generated; when the display attribute is the hidden attribute, the corresponding accompanying element is determined as an unrecognizable element, a use inclination of the unrecognizable element is identified, and when the use inclination is greater than a limit inclination, addition adjustment information is generated.

[0014] Optionally, in a possible implementation manner of the first aspect, the dynamic change of the subject element and the accompanying element in the video data according to the video recognition strategy comprises: identifying a change state of an element part corresponding to the subject element in the video data, and generating instruction adjustment information based on the change state; when the change state is independent change, a part change amplitude of the element part is identified, and when the part change amplitude is greater than a preset part amplitude, preset instruction information corresponding to the element part is called as the instruction adjustment information; when the change state is combined change, the corresponding element part is determined as a combined part, and when the combined part has an intersection, preset instruction information corresponding to the combined part is used as the instruction adjustment information.

[0015] In a second aspect of the present application, a mobile terminal edge intelligent inference system supporting multi-modal input is provided, comprising: a splicing module configured to receive a target image of a target region collected by a monitoring device in response to preposition completion information, identify a static planning area and a dynamic splicing area of each target element in the target image, and splice the target image to obtain a target recognition area of the target element; The calling module is configured to call multi-modal data of the trigger identification area based on a target amplitude of a target element in the target identification area, wherein the multi-modal data includes input data and output data; The analyzing module is configured to perform data analysis on the input data to obtain active adjustment information corresponding to the input data, and identify dynamic changes of a subject element and a companion element in the output data based on a change identification strategy to obtain trigger adjustment information of the output data. The sending module is configured to send the active adjustment information and the trigger adjustment information to an adjustment device of the target element.

[0016] In a third aspect, the present application provides an electronic device, comprising a memory, a processor and a computer program, wherein the computer program is stored in the memory, and the processor executes the computer program to perform the method of the first aspect and various possible aspects related to the first aspect.

[0017] The present application has the following advantages: 1. The present application can more comprehensively and accurately identify the demand information of the conference participants in the conference process by analyzing and processing the multi-modal data corresponding to the conference participants, thereby providing more accurate services for the conference participants and improving the experience of the conference participants.

[0018] 2. In the present application, the target identification area corresponding to the target element, i.e. the conference participant, can be accurately obtained by region identification and splicing operation when obtaining the multi-modal data corresponding to the conference participant. The multi-modal data corresponding to the target element can be more comprehensively and accurately collected through the target identification area. Specifically, the position contour of the static position element and the placement contour of the static placement element can be determined and processed to obtain a static planning area, and then the dynamic splicing area can be obtained through mirror copying and dynamic filling processing. The target identification area corresponding to the target element can be obtained by splicing the static planning area and the dynamic splicing area. A region containing the conference participant and related articles can be accurately delimited, which provides a comprehensive and accurate data basis for subsequent demand identification based on the region, and can improve the accuracy and reliability of demand identification.

[0019] 3. After the target identification area is determined, the target amplitude can be used as a screening standard to quickly locate the trigger identification area that needs to be focused on in the conference data, and the multi-modal data of the trigger identification area can be accurately called. This can comprehensively capture various demands of the conference participants in the conference process, reduce unnecessary data processing amount, and improve the efficiency and accuracy of demand analysis.

[0020] 4、The application can respectively analyze the input data and the output data when analyzing the multi-modal data, accurately obtains the active adjustment information corresponding to the input data and the trigger adjustment information corresponding to the output data through different strategies and algorithms, so that the demand of the target element in the conference process can be comprehensively reflected, the active adjustment information can be obtained by analyzing the input data, the trigger adjustment information can be obtained by analyzing the dynamic changes of the subject element and the accompanying element in the output data through the change identification strategy, so that the demand signals of the participants can be comprehensively captured, including the demand expressed through the input data and the potential demand expressed through the output data, and the accuracy and comprehensiveness of the demand identification of the participants can be improved.

[0021] 5、The application can respectively analyze the video data and the audio data when analyzing the output data, in the aspect of audio data processing, the dynamic changes of the subject element and the accompanying element in the audio data are analyzed through the audio identification strategy, the demand scene of the participants can be accurately judged according to the change of the sound characteristics, in the aspect of video data processing, the action, expression of the subject element and the change of the accompanying element in the video data are analyzed through the video identification strategy, and then the corresponding instruction adjustment information is generated, through the fine analysis of the audio and video data, the service instruction can be accurately generated according to the specific state of the participants, the accurate response to the demand of the participants is realized, and the conference experience of the participants is improved. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a flow diagram of a mobile edge intelligent push method supporting multi-modal input provided by an embodiment of the application; Figure 2 is a schematic diagram of determining the dynamic splicing area of the target element provided by an embodiment of the application; Figure 3 is a structural schematic diagram of a mobile edge intelligent push system supporting multi-modal input provided by an embodiment of the application; Figure 4 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the application clearer, the technical scheme in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0024] The technical solutions of the present application will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes may not be described in detail in some examples.

[0025] Referring to Figure 1 , a flowchart of a mobile terminal edge intelligent inference method supporting multi-modal input provided by an embodiment of the present application, Figure 1 The execution subject of the method shown can be a software and / or hardware device. The execution subject of the present application can include but is not limited to at least one of the following: user equipment, network equipment, etc. Among them, the user equipment can include but is not limited to computers, smart phones, personal digital assistants (Personal Digital Assistant, PDA) and the above-mentioned electronic devices, etc. The network equipment can include but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing, wherein cloud computing is a kind of distributed computing, which is a super virtual computer composed of a group of loosely coupled computers. This embodiment does not make any limitation. Including steps S1 to S4, as follows: S1, in response to the pre-completion information, receiving the target image collected by the monitoring device in the target area, identifying the static planning area and the dynamic splicing area of each target element in the target image and splicing to obtain the target identification area of the target element.

[0026] Among them, the pre-completion information refers to the information corresponding to the situation that all participants have completed the preparation work before starting the meeting, for example, when it is detected that all participants have placed their own water cups, fountain pens and other articles, the corresponding pre-completion information can be responded to, the monitoring device refers to the device for monitoring the personnel in the corresponding place in real time, for example, it can be a monitoring camera installed in the conference room, the target area refers to the area where the meeting is held, for example, it can be the area corresponding to the conference room, the target image refers to the image obtained by the monitoring device for image collection in the target area, the target element refers to the personnel participating in the meeting, the static planning area refers to the area corresponding to the position of the target element, the dynamic splicing area refers to the area containing all articles corresponding to the target element, such as water cups, fountain pens, etc. The target identification area refers to the area in which the demand of the target element in the meeting process can be identified.

[0027] It can be understood that whether it is an internal project seminar of an enterprise or an academic exchange meeting, it is expected to accurately identify the various needs of the participants in real time to facilitate the effective transmission of information and the rapid generation of decisions, but the existing technology can only realize the inference analysis of single modal data, and often cannot fuse and analyze the multi-modal data of the participants in the meeting process, for example, it may be difficult to process the speech data, video data, and text data of the participants in the meeting process, so it may not be possible to accurately identify the various needs of the participants in the meeting process, and thus it may not be possible to provide appropriate services for the participants. The present scheme can more comprehensively and accurately identify the demand information of the participants in the meeting process by analyzing and processing the multi-modal data corresponding to the participants, and can more accurately provide appropriate services for the participants, thereby improving the experience of the participants.

[0028] In actual application, when the monitoring device monitors that all participants participating in the meeting have completed the preparation work, for example, the water cup, pen and other articles are placed, the corresponding pre-completion information can be sent, after receiving the pre-completion information, the target image obtained by the monitoring device, for example, the monitoring camera in the meeting room, can be obtained. The target image contains a plurality of target elements, by analyzing the target image, the position of each target element in the target image, i.e., the participant, can be identified, and the corresponding region is determined as a static planning area. At the same time, the region where the article of the target element is located can be identified as a dynamic splicing area. The static planning area and the dynamic splicing area are spliced to obtain a target recognition area corresponding to each target element. The target recognition area is a basic area for subsequent demand identification, integrates the information of the target element and its related articles, and can provide image-level data support for comprehensive analysis of the demand of the target element.

[0029] In some embodiments, the step S1 of "identifying the static planning area and the dynamic splicing area of each target element in the target image and splicing to obtain the target recognition area of the target element" includes the following steps: S11, identifying the position contour of the static position element and the placement contour of the static placement element in the target image.

[0030] The static position element refers to the chair on which the participant sits, the position contour refers to the contour corresponding to the static position element, and the static placement element can be a conference table, and the placement contour refers to the contour corresponding to the static placement element.

[0031] Specifically, the image recognition technology can be used to preliminarily analyze the target image. The static position element refers to the chair on which the participant sits. The position and contour of the chair can reflect the basic position information of the participant in the conference space. The static placement element can be a conference table. The placement contour corresponding to the conference table can define the main planar region of the conference activity and is closely related to the position of the participant. Through edge detection, feature extraction and other algorithms, the boundary lines of the chair in the image can be identified to form the corresponding position contour. At the same time, the edge of the conference table can be accurately positioned to obtain the corresponding placement contour. These contour information is the basis for subsequent region division and processing and provides a key reference for accurately determining the region where the participant is located.

[0032] S12, determining the position contour closest to each placement edge length in the placement contour as the processing contour of each placement edge length.

[0033] After obtaining the placement contour corresponding to the static placement element and the position contour corresponding to the static position element, the spatial relationship between the two can be analyzed. For each edge length of the placement contour (conference table contour), the distance between it and each position contour (chair contour) can be calculated to find the position contour closest to the edge length, which is determined as the processing contour corresponding to the placement edge length. This step can clearly define the spatial correspondence between the conference table and the chair, provide a basis for subsequent processing of the region based on this relationship, and ensure that the generated region can closely surround the participant and his / her surrounding environment, so that the divided region is more consistent with the activity range of the participant in the actual scene.

[0034] In the formula, the placement edge length refers to the edge length corresponding to the placement contour, and the processing contour refers to the position contour closest to each placement edge length.

[0035] S13, processing the fixed edge length and the extended edge length in the processing contour to obtain the static planning region of the target element.

[0036] After determining the processing contour, the edge lengths of the processing contour can be further classified and processed. The fixed edge length refers to the relatively stable edge reflecting the inherent shape and position of the chair. The extended edge length is an edge that may need to be extended or adjusted according to factors such as the position relationship with the conference table. Through specific algorithms and rules, the fixed edge length is maintained in its original form, and the extended edge length is appropriately extended. For example, according to the relative position of the conference table and the chair, the edge of the chair perpendicular to the conference table is reasonably extended to better cover the activity range that the participant may involve in the conference process. After processing the fixed edge length and the extended edge length, a more accurate region, i.e., the static planning region of the target element (participant), can be obtained. This region can more accurately define the basic position range of the participant in the conference space.

[0037] The fixed side length refers to a relatively fixed side in the processing contour, and the extended side length refers to a side that can be extended in the processing contour.

[0038] On the basis of the above embodiment, the specific implementation mode of step S13 can be: S131, the side length in the processing contour parallel to the corresponding placement side length is taken as the fixed side length, and the remaining side lengths are taken as the extended side lengths.

[0039] Specifically, each side of the processing contour can be analyzed one by one to determine whether it is parallel to a side length of the conference table placement contour. If a side is parallel to a conference table placement side length, this side can be identified as a fixed side length. The remaining side lengths in the processing contour that are not parallel to the conference table placement side length can be defined as extended side lengths. Through this parallel relationship-based side length classification mode, the processing direction of different side lengths can be determined, and preparation for constructing an accurate static planning area is made.

[0040] S132, the fixed side length closest to the corresponding placement side length is deleted, and the extended side length is extended to one side of the corresponding placement side length until it intersects with the placement side length, to obtain the static planning area of the target element.

[0041] Specifically, after the fixed side length is determined, the fixed side length closest to the conference table placement side length can be deleted. Since the fixed side length closest to the conference table limits the expansion of the area in the direction of the conference table to some extent, and the conference table is an area where the participants frequently contact and move during the meeting, deleting this side can create space for the extension of the extended side length, so that the delineated area can better cover the activity range of the participants around the conference table. Then, the extended side length can be extended, and each extended side can be extended to one side of the corresponding conference table placement side. The extension process continues until the extended side intersects with the placement side of the conference table. In this way, it can be ensured that the delineated area can closely fit the spatial relationship between the conference table and the chair, and fully cover the activity area that the participants may involve during the meeting. When all the extended sides are extended to intersect with the placement side of the conference table, the final area formed is the static planning area of the target element (the participant), which accurately defines the basic position range of the participant in the conference space, considering not only the inherent shape and position of the chair, but also the relative position relationship between the conference table and the chair, providing an accurate position basis for subsequent identification and analysis of the participant's needs based on the area.

[0042] S14, mirroring and copying the corresponding static planning area based on the placement side length to obtain a copied planning area.

[0043] Specifically, with the placement side length of the conference table as a reference, the generated static planning area can be mirror copied. In an actual conference scene, mirror copying can be performed with the sides of the conference table as the axis of symmetry, which can simulate the area where the participants may have items or activities on the conference table, providing a basic framework for subsequent consideration of the area where all items around the participants are located, so that the generated area can more comprehensively cover the item placement and activity range that the participants may be involved in. The copied planning area refers to the area obtained by mirror copying the static planning area according to the placement side length.

[0044] S15, performing dynamic filling processing on the copied planning area according to the accompanying contour of the accompanying placement element in the placement contour, to obtain a dynamic splicing area of the target element.

[0045] In actual applications, the accompanying placement elements placed on the conference table by the participants can be located on the boundary of the copied planning area or outside the copied planning area. The accompanying placement elements can be various items placed on the conference table or around the participants, such as a cup, a pen, a notebook, etc. The accompanying contour refers to the boundary shape of these items in the image. According to the accompanying contour of the accompanying placement element, the copied planning area can be further processed. For the part of the copied planning area that does not cover the accompanying placement element, dynamic filling, expansion, or the like can be performed to fill in the gaps. For example, if there is a cup outside the copied planning area, the copied planning area can be expanded according to the contour of the cup to include the area where the cup is located. After this dynamic filling processing, a dynamic splicing area corresponding to the target element can be obtained, which completely contains the area where all items of the participant are located, providing more abundant image information for comprehensive analysis of the participant's needs.

[0046] The accompanying placement element refers to various items placed on the conference table or around the participants, such as a cup, a pen, a notebook, etc. The accompanying contour refers to the contour corresponding to the accompanying placement element.

[0047] On the basis of the above embodiment, the specific implementation of step S15 can be: S151, identifying the accompanying contour of the accompanying placement element in the placement contour, taking the accompanying contour that has an intersection with the copied planning area as a merging contour, and taking the remaining accompanying contour as a to-be-selected contour.

[0048] Specifically, image recognition technology can be used to recognize various types of accompanying placement elements (such as a cup, a pen, a notebook, etc.) in the placement contour corresponding to the conference table, and extract the boundary shape of these objects in the target image, that is, the accompanying contour. Each extracted accompanying contour is subjected to spatial relationship judgment with the previously generated replication planning area. The basis for the judgment is whether there is an intersection between the two areas. If there is an overlapping part between the accompanying contour and the replication planning area in the image space, the accompanying contour is marked as a merged contour. The objects corresponding to these merged contours are already partially in the replication planning area, and they need to be completely included in the area in the future. For those accompanying contours that do not intersect with the replication planning area, they are classified as to-be-selected contours. These objects are completely outside the replication planning area and need to be further judged whether to be included in the extended range.

[0049] Among them, the merged contour refers to the accompanying contour that intersects with the replication planning area, and the to-be-selected contour refers to the accompanying contour that does not intersect with the replication planning area.

[0050] S152, according to the union set of the merged contour and the corresponding replication planning area, a first combined contour is obtained.

[0051] After determining the merged contour, the merged contour and the corresponding replication planning area can be subjected to a union operation in set operation. The purpose of the union operation is to completely integrate the area of the accompanying placement element that is partially located in the replication planning area into a new area. In this way, the replication planning area and the area of the object represented by the accompanying contour that intersects with it can be fused, the boundary between the areas can be eliminated, and a more complete area contour containing part of the accompanying placement element, that is, the first combined contour, can be formed. Compared with the original replication planning area, the first combined contour has been preliminarily expanded in scope and is more in line with the actual distribution of the objects used by the participants, which can prepare for further expanding the area to include more objects.

[0052] Among them, the first combined contour refers to the contour obtained by performing a union operation on the merged contour and the corresponding replication planning area.

[0053] S153, coordinate processing is performed on the target image to obtain the straight-line distance from the center point of the to-be-selected contour to the first combined contour, and the to-be-selected contour with a straight-line distance less than a preset screening distance is selected as a selected contour of the corresponding first combined contour.

[0054] In order to further determine whether the accompanied placement element outside the replication planning area is included in the dynamic splicing area, the target image can be subjected to coordinate processing, a coordinate system of the image is established, and a coordinate value is assigned to each pixel point in the image, so that the spatial distance between different regions can be accurately calculated. For each to-be-selected contour (i.e., an accompanied contour without intersection with the replication planning area), the coordinates of the center point thereof can be obtained, and then the shortest straight line distance from the center point to each point on the first combined contour is calculated. Whether the to-be-selected contour is included in the extended range is determined by comparing the size relationship between the straight line distance and a preset screening distance. The preset screening distance is a threshold value preset according to the possible distribution range of the articles of the participants in the actual conference scene. If the straight line distance from the center point of the to-be-selected contour to the first combined contour is less than the preset screening distance, it is indicated that the article is outside the replication planning area, but is close to the integrated region, and is likely to be an article used by the participants in the conference process. Therefore, the to-be-selected contour is taken as a selected contour. In this way, the region integrity can be ensured, and the region is not excessively expanded, so that the dynamic splicing area can cover the important article region and does not include too many irrelevant regions, and the accuracy and effectiveness of region construction are improved.

[0055] The straight line distance refers to the vertical distance between the center point of the to-be-selected contour and the first combined contour, the preset screening distance refers to a distance preset for measuring the size of the straight line distance, and the selected contour refers to the to-be-selected contour with the straight line distance less than the preset screening distance.

[0056] In S154, the coordinate extreme values of the selected contour and the first combined contour are obtained, a reconstruction rectangle is constructed based on the coordinate extreme values, and the reconstruction rectangle is taken as the dynamic splicing area of the target element.

[0057] After the selected contour is determined, the selected contour and the first combined contour can be analyzed and processed. The coordinate extreme values of the contours in the image coordinate system, i.e., the minimum value and the maximum value of the horizontal coordinates and the minimum value and the maximum value of the vertical coordinates of all contour points, can be obtained. The coordinate extreme values can determine a minimum rectangular range capable of completely containing the selected contour and the first combined contour. According to the coordinate extreme values, a reconstruction rectangle corresponding to the minimum rectangular range can be constructed. The reconstruction rectangle is the final dynamic splicing area of the target element, and can completely contain the region where all articles of the participants are located.

[0058] Referring to Figure 2 A schematic diagram for determining the dynamic splicing area of the target element provided by the embodiment of the application is shown in Figure 2 After the first combined contour is obtained, when the straight line distance between the to-be-selected contour and the first combined contour on the placement contour is less than the preset screening distance, as shown in Figure 2The corresponding to-be-selected contour can be determined as a selected contour according to the circular contour in the image, and a minimum rectangle containing the selected contour and the first combined contour can be constructed according to coordinate extreme values of the selected contour and the first combined contour, that is, a reconstruction rectangle, as shown in Figure 2 The area corresponding to the reconstruction rectangle can be determined as the dynamic splicing area.

[0059] By constructing the reconstruction rectangle, subsequent unified management and analysis of the area are facilitated, whether it is to retrieve image data of the area or to perform fusion analysis of multi-modal data based on the area, which is more convenient and efficient. The formation of the dynamic splicing area provides a comprehensive and regular image information basis for subsequent accurate identification of participant needs according to the target identification area, and can more accurately analyze the relationship between the participant and the surrounding objects, so that the participant's needs during the meeting can be better understood.

[0060] The coordinate extreme values include maximum and minimum horizontal coordinate values and maximum and minimum vertical coordinate values corresponding to the selected contour and the first combined contour, and the reconstruction rectangle refers to a rectangle constructed according to the coordinate extreme values.

[0061] On the basis of the above embodiment, when the straight-line distances between the to-be-selected contour and the plurality of first combined contours are all less than the preset screening distance, that is, there may be an object between at least two participants, which may be shared by the plurality of participants, for example, two participants may share a notebook computer. In this case, the area corresponding to the shared object can be separately delimited, and then the separately delimited area and the dynamic splicing area corresponding to each participant can be connected by a marking line to indicate that the plurality of participants share the corresponding object. Specifically, the contour edge closest to each first combined contour in the to-be-selected contour can be determined as a to-be-selected contour edge. For example, when contour edge 1 and the first combined contour 1 are very close, and contour edge 3 and the first combined contour 2 are very close, contour edges 1 and 3 can be determined as to-be-selected contour edges, and then the center points of the to-be-selected contour edges can be connected to the corresponding first combined contours, for example, the center point of contour edge 1 can be connected to the first combined contour 1, and the center point of contour edge 3 can be connected to the first combined contour 3, indicating that the object corresponding to the to-be-selected contour is an object shared by the participant corresponding to the first combined contour 1 and the participant corresponding to the first combined contour 3.

[0062] S16, splicing the static planning area and the dynamic splicing area to obtain a target identification area of the target element.

[0063] After obtaining the static planning area and the dynamic stitching area respectively, the two areas can be integrated. For example, the static planning area (representing the position of the participant) and the dynamic stitching area (containing the area of the participant's belongings) can be seamlessly stitched together by image fusion technology to form a complete area, i.e., a target recognition area of the target element. The target recognition area integrates the position information of the participant and the surrounding information of the belongings, and can become a basic area for subsequent identification and analysis of the participant's needs in the conference process. Based on the target recognition area, further multi-modal data can be retrieved to analyze and judge the participant's needs more deeply and comprehensively, laying a solid foundation for providing adaptive services.

[0064] Through the above implementation, the possible activity area of the participant can be accurately positioned, the irrelevant area information is reduced to interfere with the demand judgment, and the pertinence in demand identification is improved.

[0065] S2, based on the target amplitude of the target element in the target recognition area, the corresponding target recognition area is taken as a trigger recognition area, and multi-modal data of the trigger recognition area is retrieved, the multi-modal data including input data and output data.

[0066] Among them, the target amplitude refers to the amplitude corresponding to the action of the target element in the conference process, the trigger recognition area refers to the target recognition area corresponding to the target element with a larger target amplitude, and the multi-modal data refers to data that can infer the demand information of the target element, including input data actively input by the target element and output data such as audio data and video data collected corresponding to the target element.

[0067] In the conference scenario, if the participant has a demand, he / she will often express it through a larger amplitude action, for example, when the participant cannot see the PPT content, he / she may perform actions such as leaning forward and stretching the neck. On the contrary, the participant with a smaller action amplitude can be considered to have no need for external assistance in the conference process under normal circumstances. This way of using action amplitude as a screening standard can quickly locate the area that really needs attention in a large amount of conference data, greatly reducing unnecessary data processing.

[0068] Specifically, first, the action of the target element in each target recognition area can be monitored in real time, for example, the action of the target element can be analyzed frame by frame, and the value of the target amplitude can be quantified by calculating the displacement and angle change of the target element within a certain time, for example, by identifying the moving distance of the participant's head or torso and the angle of the raised arm, and combining a preset amplitude threshold to determine whether the target amplitude is large. When the target amplitude of a certain target element exceeds the preset threshold, the target recognition area corresponding to the target element can be marked as a trigger recognition area.

[0069] After the trigger recognition area is determined, the multi-modal data of the area can be called, and the multi-modal data covers multiple dimensions capable of inferring the demand information of the target element. The input data is derived from the content input by the target element through the device, such as sending a text message using a conference terminal, operating a remote controller to issue an instruction, etc. These data can directly reflect the demand explicitly expressed by the target element. The output data is mainly audio data and video data collected by monitoring the device, such as the speech content, facial expressions, and body movements of the participants, etc. These data are passively collected, but often contain the potential and not yet explicitly expressed demand information of the target element. In the data calling process, the corresponding data segments can be accurately positioned and extracted according to the range of the trigger recognition area. For video data, the video stream of the trigger recognition area in a specific time period can be obtained. For audio data, the sound signals emitted by the target element in the area are filtered out. For input data, all information input by the target element through the related device in the time period is searched. In this way, the multi-modal data called is both comprehensive and accurate, providing rich and effective data support for the subsequent in-depth analysis of the demand of the target element.

[0070] In some embodiments, step S2 includes S21, specifically as follows: S21, identifying the target amplitude of the target element in the target recognition area, and determining that the target amplitude is greater than a preset amplitude to take the corresponding target recognition area as the trigger recognition area.

[0071] Specifically, real-time action monitoring can be performed on the target elements (attendees) in each target recognition area. For example, advanced image recognition and analysis techniques, such as key point detection algorithms in the field of computer vision, can be used to analyze the actions of the target elements frame by frame. The head, torso, limbs, and other key parts of the attendees are detected as the detection objects. By calculating the displacement and angle change of these parts within a certain period of time, the action amplitude values of the target elements are quantified. For example, by tracking the movement distance of the head and the angle of the raised arm, and combining a preset calculation formula, the action is converted into a specific target amplitude value. In a normal conference scenario, the action amplitude of the attendees is usually within a certain reasonable range. For example, when the attendees are seriously listening, taking notes, or performing routine body relaxation, the action amplitude of the key parts such as the head, torso, and limbs is relatively small and has a certain regularity. Through monitoring and data analysis of the actions of the attendees in a large number of actual conference scenarios, the amplitude range of these natural actions is statistically obtained, and a preset amplitude threshold, i.e., a preset amplitude, is set. When the target amplitude of the recognized target element is greater than the preset amplitude, it can be considered that the action of the target element has exceeded the range of normal natural activities, which may indicate that the target element has encountered a problem or has a certain need for external assistance during the conference. Therefore, the target recognition area corresponding to the target element can be marked as a trigger recognition area. The preset amplitude refers to a threshold for measuring the action amplitude of the target element, which is set in advance based on the statistical analysis results of the natural activity amplitude of the body parts of the attendees in a normal conference scenario.

[0072] Through the above implementation, the area with potential needs can be quickly locked. Compared with comprehensive processing of the entire conference scene data, the real-time method can reduce unnecessary data processing amount and improve data processing efficiency.

[0073] S3, performing data analysis on the input data to obtain active adjustment information corresponding to the input data, identifying the dynamic changes of the subject elements and the accompanying elements in the output data based on a change identification strategy, and obtaining trigger adjustment information of the output data.

[0074] The active adjustment information refers to information corresponding to the input data and obtained after analyzing the input data, which is used to provide corresponding assistance for the corresponding target element. The change identification strategy refers to a method and rule for analyzing the dynamic changes of the subject element and the accompanying element in the output data. The subject element refers to an object that needs to be focused on in the output data, and the specific meaning varies depending on the type of output data. In audio data, the subject element is usually the sound corresponding to the target element (meeting participant), i.e., the sound information such as speech, questions, etc. emitted by the meeting participant. In video data, the subject element is generally the target element itself, i.e., the image of the meeting participant, including its actions, expressions, postures, etc. The accompanying element refers to an element associated with the subject element and plays an auxiliary analysis role in the output data. It also varies depending on the type of output data. In audio data, the accompanying element is other background sound in addition to the sound of the target element, such as environmental noise in the meeting room, running sound of other equipment, etc. These background sounds may affect the clarity and analysis of the subject element (meeting participant sound). In video data, the accompanying element can be the personal items of the target element, such as the files, water cup, notebook computer, etc. placed on the table by the meeting participant. The status and changes of these items may also reflect the needs of the meeting participant and trigger adjustment information. The trigger adjustment information refers to information obtained after analyzing the dynamic changes of the subject element and the accompanying element in the output data using the change identification strategy, which is used to trigger adjustments related to the meeting.

[0075] After obtaining the multi-modal data, the input data and the output data can be processed respectively. For the input data, existing data analysis techniques can be used to convert it into active adjustment information corresponding to the input data. For example, if the meeting participant inputs an instruction to adjust the screen brightness through the device, the instruction will be converted into specific active adjustment information after analysis, indicating that the screen brightness needs to be adjusted.

[0076] For output data, a change identification strategy can be used for analysis. The output data contains subject elements and accompanying elements, which can be different in different output data. For example, for audio data, the subject element can be the sound corresponding to the target element, and the accompanying element can be other background sound other than the sound of the target element. For video data, the subject element can be the target element itself, and the accompanying element can be the personal belongings of the target element. Through the change identification strategy, the dynamic changes of the subject elements and the accompanying elements are identified. For example, when the participant frequently raises his hand and has a puzzled expression, through the identification and analysis of these dynamic changes, the trigger adjustment information of the output data is obtained. It can be judged that the participant has doubts about the meeting content and needs staff to answer. Through the processing of the input data and the output data, the active adjustment information and the trigger adjustment information can be obtained. These two types of information can fully reflect the needs of the target element in the meeting process.

[0077] In some embodiments, the step of "identifying the dynamic changes of the subject elements and the accompanying elements in the output data based on the change identification strategy, and obtaining the trigger adjustment information of the output data" in step S3 comprises the following steps: S31, obtaining output data, wherein the output data comprises video data and audio data.

[0078] Specifically, the output data corresponding to the determined trigger identification area can be obtained, including audio data and video data related to the area. These data are continuously collected and stored by the monitoring device during the meeting process, and contain various behavior information of the target element (participant) in the trigger identification area. For example, if the trigger identification area is the area where a participant is located in the meeting room, the video stream corresponding to the area can be extracted from the stored meeting video data, and the sound signal generated in the area can be extracted from the audio recording. Through accurate acquisition of these data, raw materials are provided for subsequent demand analysis based on audio and video data, which is the starting point of the entire output data analysis process.

[0079] Among them, the video data refers to the dynamic image data related to the trigger identification area continuously collected and stored by the monitoring device during the meeting process, and the audio data refers to the sound data related to the trigger identification area continuously collected and stored by the monitoring device during the meeting process.

[0080] S32, based on the audio identification strategy, the dynamic changes of the subject elements and the accompanying elements in the audio data are obtained, and the trigger adjustment information of the audio data is obtained.

[0081] Specifically, the audio recognition strategy is an analysis method and rule formulated according to the characteristics of audio data. In audio data, the main element is the sound emitted by the target element (conference participant), such as speech content, tone of voice, questions, etc. The accompanying element is other background sound other than the sound of the target element, such as conference room environmental noise, device operation sound, etc. The characteristics of the sound, such as volume, speed, and tone, can be analyzed by a speech recognition algorithm. Different changes in sound characteristics correspond to different demand scenarios. By quantitatively analyzing these characteristics, the conference process can be accurately positioned to adjust the link, and more personalized services can be provided to conference participants. For example, a sudden increase in volume may indicate that the conference participant is coughing, and a paper towel may be needed. In addition, changes in background sound, such as sudden noise, may affect communication between conference participants, and it may be necessary to adjust the volume of the audio equipment or optimize the conference environment. By comprehensively analyzing the dynamic changes of the main element and the accompanying element in the audio data, the audio recognition strategy can obtain the trigger adjustment information corresponding to the audio data, such as whether the conference participant needs to repeat the explanation of the content or adjust the settings of the conference audio equipment.

[0082] Among them, the audio recognition strategy refers to a strategy for determining trigger adjustment information by recognizing audio data.

[0083] On the basis of the above embodiment, the specific implementation mode of step S32 can be: S321, identifying the main sound wave in the audio data as the main element and the remaining sound wave as the accompanying element based on the timbre.

[0084] Specifically, the timbre recognition technology can be used to preliminarily divide the audio data. The timbre is the characteristic of the sound. Different sound sources (conference participants, devices in the conference environment, and other background sound sources) have unique timbre characteristics. By analyzing each sound wave in the audio data using a pre-trained timbre recognition model, the model is trained based on a large number of audio samples containing different human voices and environmental sounds, and can accurately identify the sound wave emitted by the conference participant. The sound wave representing the voice of the conference participant is marked as the main sound wave and determined as the main element, which contains key information such as conference participant speech and questions. The remaining sound waves of non-conference participant voices, such as conference room air conditioner operation sound and table and chair movement sound, are classified as accompanying elements. This division based on timbre lays the foundation for subsequent analysis of the dynamic changes of the main element and the accompanying element, so that it can focus on sound information directly related to the needs of conference participants, and take into account background sounds that may affect conference communication. Among them, the main sound wave refers to the sound wave corresponding to the conference participant, the main element in the audio data refers to the main sound wave, and the accompanying element refers to other interfering sound waves in the audio data other than the main sound wave.

[0085] S322, extract a subject amplitude of a subject element in the audio data, and determine that the subject amplitude is greater than a preset subject amplitude, and take corresponding sound waves as subject change sound waves, and obtain a subject duration of each of the subject change sound waves.

[0086] Specifically, after obtaining the subject element, the subject element can be further analyzed, and the amplitude information corresponding to the subject sound waves, i.e., the subject amplitude, can be extracted. The amplitude can reflect the strength of the sound, and the subject amplitude is the strength value of the sound of the participant. A preset subject amplitude can be set in advance. The threshold is determined according to the normal conference communication sound intensity range and the actual business demand. When it is detected that the amplitude of the subject sound waves is greater than the preset subject amplitude, it indicates that the sound of the participant has changed obviously at this time. These sound waves can be marked as subject change sound waves. For example, the sudden increase in the volume of the participant speaking will trigger this marking. At the same time, the duration of each subject change sound wave, i.e., the subject duration, can be recorded. By obtaining the subject change sound waves and the subject duration thereof, the specific segment and the duration of the change in the sound intensity of the participant can be captured, and key data for subsequent analysis of the mood and intention of the participant can be provided.

[0087] Among them, the subject amplitude refers to the amplitude corresponding to the subject sound waves, the preset subject amplitude refers to the amplitude value set in advance for measuring the size of the subject amplitude, the subject change sound wave refers to the subject sound wave with an amplitude greater than the preset subject amplitude, and the subject duration refers to the duration corresponding to the subject change sound wave.

[0088] S323, determine that the subject duration is less than a preset duration, take corresponding subject change sound waves as instantaneous subject sound waves, and retrieve preset transmission information based on the instantaneous subject sound waves.

[0089] After obtaining the subject change sound waves and the subject duration thereof, the subject duration can be compared with a preset duration. The preset duration is a time threshold set according to common short-time sound change scenarios (such as sudden exclamation and short question). If the subject duration is less than the preset duration, it indicates that the subject change sound wave belongs to the change in the sound intensity in a short time, and it can be defined as an instantaneous subject sound wave. For this kind of instantaneous subject sound wave, corresponding preset transmission information can be retrieved from a preset information library. The preset information library stores common demand judgments for different instantaneous sound change scenarios. For example, when a short and high-intensity sound is detected, the preset transmission information "the participant may have an urgent problem to be quickly responded" can be retrieved. This step can quickly respond to the abnormal change in the sound of the participant in a short time, and timely prompt the relevant personnel to pay attention to the possible demand.

[0090] The preset time length is a time threshold value pre-set according to a common short-time sound change scene (for example, a sudden exclamation, a brief question, etc.), the instantaneous subject sound wave is a subject change sound wave with a subject time length less than the preset time length, and the preset delivery information is common demand judgment information pre-set in a preset information library for different instantaneous sound change scenes.

[0091] In S324, when the subject time length is greater than the preset time length, the corresponding subject change sound wave is determined as a continuous subject sound wave, the text information corresponding to the continuous subject sound wave is extracted, and the labeling adjustment information is generated based on the text information.

[0092] Specifically, when the subject time length is greater than the preset time length, it indicates that the subject change sound wave is a sound intensity change with a relatively long duration, which can be identified as a continuous subject sound wave. It can be considered that the sound wave with a relatively long duration and a large intensity may be important content in the meeting. The continuous subject sound wave can be converted into text information by using a voice recognition technology, for example, the content of the continuous speech of the meeting participant can be converted into readable text. Then, the semantic analysis is performed on the text information to understand the content and intention expressed by the meeting participant. According to the analysis result, the corresponding labeling adjustment information can be generated to highlight the corresponding meeting content, so as to facilitate the subsequent personnel to check the highlighted meeting content.

[0093] The continuous subject sound wave is a subject change sound wave with a subject time length greater than a preset time length, the text information is readable text content converted from the continuous subject sound wave by using a voice recognition technology, and the labeling adjustment information is specific information generated by understanding the content and intention expressed by the meeting participant through semantic analysis of the text information converted from the continuous subject sound wave, and is used to indicate a highlighting operation on the corresponding meeting content. The labeling adjustment information can guide the subsequent relevant personnel (such as a meeting recorder, a meeting participant, etc.) to quickly locate and check the key content in the meeting, and facilitate the arrangement, review and use of the key information in the meeting.

[0094] In S325, the accompanying amplitude of the accompanying element in the audio data is extracted, and when the accompanying amplitude is greater than a preset accompanying amplitude, the corresponding sound wave is determined as an accompanying change sound wave, and the accompanying time length of each accompanying change sound wave is obtained.

[0095] Specifically, after analyzing the main elements, the accompanying elements can be analyzed accordingly, and the amplitude of the accompanying sound wave, i.e., the accompanying amplitude, can be extracted, which can represent the strength of the background sound in the conference environment. Similarly, a preset accompanying amplitude can be set to measure whether the background sound is abnormally enhanced. When the accompanying amplitude is greater than the preset accompanying amplitude, it means that the intensity of the background sound in the conference environment exceeds the normal range. These sound waves can be marked as accompanying change sound waves, and their duration, i.e., accompanying duration, is recorded. For example, a larger equipment failure sound suddenly occurs in the conference room. The corresponding equipment failure sound can be marked as an accompanying change sound wave and the duration is recorded. This step can timely discover sound changes in the conference environment that may interfere with normal communication and provide data support for subsequent measures.

[0096] wherein the accompanying amplitude refers to the amplitude of the accompanying sound wave, the preset accompanying amplitude refers to a set value for measuring whether the background sound is abnormally enhanced, the accompanying change sound wave refers to the sound wave marked when the accompanying amplitude is greater than the preset accompanying amplitude, and the accompanying duration refers to the duration of the accompanying change sound wave.

[0097] S326, determining that the accompanying duration is less than the preset duration, regarding the corresponding accompanying change sound wave as a transient accompanying sound wave, and based on the transient accompanying sound wave, retrieving disposal adjustment information.

[0098] Specifically, after obtaining the accompanying duration, the accompanying duration can be compared with the preset duration set in advance. If the accompanying duration is less than the preset duration, it means that the accompanying change sound wave is an abnormal background sound in a short time, which is defined as a transient accompanying sound wave. For the transient accompanying sound wave, corresponding disposal adjustment information can be retrieved from the preset disposal scheme library. For example, if the short-term noise is caused by a short-term equipment failure, the disposal adjustment information of "checking the related equipment to confirm whether it needs to be repaired" can be retrieved. By quickly responding to the transient background sound change, the sudden environmental problem that may affect the conference can be timely processed, and the normal order of the conference can be maintained. The transient accompanying sound wave refers to the accompanying change sound wave with a duration less than the preset duration, and the disposal adjustment information refers to specific information formulated for the transient accompanying sound wave to guide the adoption of corresponding processing measures.

[0099] S327, determining that the accompanying duration is greater than the preset duration, regarding the corresponding accompanying change sound wave as a continuous accompanying sound wave, generating an elimination sound wave opposite to the continuous accompanying sound wave, and taking the elimination sound wave as elimination adjustment information.

[0100] Specifically, when the accompanying time is greater than the preset time, it can be considered that the background sound is in an abnormal state for a long time, and it can be determined as a continuous accompanying sound wave. In order to eliminate the influence of the long-time abnormal background sound on the conference communication, an audio processing technology can be used to generate a cancellation sound wave opposite to the continuous accompanying sound wave in frequency, amplitude and other characteristics. For example, for continuous low-frequency noise, a low-frequency sound wave opposite in phase is generated to cancel it out. The generated cancellation sound wave is used as cancellation adjustment information, which is sent to the related audio equipment. The corresponding noise reduction program can be started to reduce the interference of background noise on the conference and create a clearer communication environment for the participants.

[0101] Among them, the continuous accompanying sound wave refers to an accompanying change sound wave with an accompanying time greater than the preset time, that is, the background sound is in an abnormal state for a long time. The cancellation sound wave refers to a sound wave generated by using an audio processing technology, which is opposite to the continuous accompanying sound wave in frequency, amplitude and other characteristics, and is used to cancel the continuous accompanying sound wave to reduce its influence. The cancellation adjustment information refers to the generated cancellation sound wave.

[0102] S33, obtaining the trigger adjustment information of the video data according to the dynamic changes of the subject element and the accompanying element in the video data according to the video recognition strategy.

[0103] Specifically, the video recognition strategy is an analysis strategy designed for video data. In the video data, the subject element is the target element (conference participant) itself, including its actions, expressions, postures and other external manifestations. The accompanying element is the personal belongings of the conference participant, such as the file and water cup on the table, etc. Computer vision technology can be used to analyze the video data frame by frame. For example, the human body posture recognition algorithm can be used to detect the body movements of the conference participant, such as frequently raising hands, which may indicate that the participant wants to speak, and leaning forward with a frown, which may indicate that the participant cannot see the screen content. The expression recognition algorithm can be used to analyze the facial expressions of the conference participant, such as a confused expression, which may indicate that the participant does not understand the conference content and needs to be reminded to explain the relevant content. At the same time, changes in accompanying elements can be observed, such as the conference participant looking at the water cup several times, which may mean that the participant needs to replenish water, and the file being frequently flipped, which may mean that the participant is looking for relevant materials. At this time, it may be necessary to provide clearer conference material display. By comprehensively considering the dynamic changes of the subject element and the accompanying element in the video data, and reasoning and judging according to the video recognition strategy, the trigger adjustment information of the video data can be obtained, such as whether the projection picture needs to be adjusted, whether drinks need to be provided to the conference participants, etc., so as to accurately capture and respond to the potential needs of the conference participants.

[0104] Among them, the video recognition strategy refers to a strategy for determining the trigger adjustment information corresponding to the target element by analyzing the video data.

[0105] On the basis of the above embodiment, the specific implementation mode of step S33 can be: S331, identify the target element in the video data as the subject element, and the accompanying placement element as the accompanying element.

[0106] Specifically, the video data can be parsed, and the target detection algorithm can be used to accurately locate the conference participants (target elements) in the video screen, and determine them as the subject elements. These algorithms can quickly and accurately identify the conference participants in the screen and outline their contours through training on a large number of image samples containing human features. At the same time, the accompanying placement elements in the target recognition area corresponding to the target elements in the video screen, such as files, cups, laptops, etc. located in the target recognition area, can be identified as accompanying elements. In the identification process, the algorithm can analyze the shape, color, texture and other characteristics of the objects, and compare them with the pre-stored object feature library, so as to realize accurate identification of various objects. By clearly distinguishing the subject elements and the accompanying elements, the foundation is laid for subsequent analysis of the dynamic changes of the two, so that the visual information related to the needs of the conference participants can be captured.

[0107] Among them, the subject element here refers to the target element in the video data, and the accompanying element refers to the accompanying placement element in the video data.

[0108] S332, identify the change state of the element part corresponding to the subject element in the video data, and generate instruction adjustment information based on the change state.

[0109] Specifically, the action, posture and other change states of each part of the body of the conference participants, i.e. the element part, in the video data can be monitored in real time. Human body posture recognition algorithms such as OpenPose can be used. This algorithm can detect multiple key joints of the human body. By analyzing the position changes and relative relationships of these joints, the body movements of the conference participants can be determined, such as raising hands, bending over, leaning forward, etc. At the same time, in combination with the expression recognition algorithm, the facial expression changes of the conference participants can be analyzed to identify expressions such as smiling, frowning, confusion, etc. For example, if it is identified that the conference participant leans forward and concentrates his gaze in the direction of the projection screen for a long time, it is presumed that he may not be able to see the screen content, and the instruction adjustment information "adjust the clarity of the projection picture or enlarge the font" can be generated. Through accurate identification and analysis of the change state of each part of the subject element, the body language and expressions of the conference participants can be converted into specific service instructions, and timely response to the explicit needs of the conference participants can be realized.

[0110] Among them, the element part refers to each part of the body of the subject element (such as the conference participant) in the video data, the change state refers to the changes in the action, posture and other aspects of each element part of the body of the conference participant, and the instruction adjustment information refers to the specific information generated for indicating the corresponding adjustment and operation according to the identification and analysis results of the change state of each element part of the conference participant.

[0111] On the basis of the above-mentioned embodiments, the specific implementation mode of step S332 can be: S3321, identifying the change state of the element part corresponding to the subject element in the video data, the change state including independent change and combined change.

[0112] Specifically, computer vision algorithms can be used to analyze video data frame by frame, monitor the changes of each part of the participants' body (element part) in real time, accurately capture the changes of each element part through human pose recognition algorithm (such as OpenPose) and expression recognition algorithm, and classify these changes. Independent change refers to the change of a single element part, such as only arm lifting, head turning, etc. Combined change refers to the change of multiple element parts, such as hand covering mouth when yawning, or body leaning forward while frowning and looking at the screen, etc. Through analysis and judgment of the changes of element parts, independent change and combined change can be distinguished, which can lay a foundation for subsequent targeted analysis strategies.

[0113] Among them, the change state refers to the change of the element part corresponding to the subject element in the video data in terms of action, posture and expression, the independent change refers to the change of a single element part in the subject element, which is only the action, posture change or expression change made by a single body part, and the combined change refers to the change of multiple element parts in the subject element, which is the combined presentation of action, posture or expression of multiple body parts.

[0114] S3322, when the change state is determined to be independent change, identifying the part change amplitude of the element part, and when the part change amplitude is determined to be greater than the preset part amplitude, retrieving the corresponding preset instruction information as instruction adjustment information based on the element part.

[0115] When the change state is determined to be independent change, the change amplitude of the element part can be further quantitatively analyzed. By calculating the displacement and angle change of the element part within a certain time, the part change amplitude can be obtained. Then, a preset part amplitude can be set in advance. The threshold value is obtained based on statistical analysis of the natural activity amplitude of the body part of the participant in the normal conference scenario. When the part change amplitude of the element part is greater than the preset part amplitude, it indicates that the action may have a specific demand intention. At this time, the corresponding preset instruction information can be retrieved from the preset instruction information library according to the element part, for example, if it is detected that the participant's head is tilted back at a large angle and maintained for a long time, it can be judged that the participant may feel tired. The preset instruction information "provide a cup of coffee for the participant" can be retrieved from the instruction information library as the instruction adjustment information to realize the timely response to the demand of the participant.

[0116] The part change amplitude refers to the degree of change of the action and posture of the element part of the subject element (such as a participant) in the video data, the preset part amplitude refers to the amplitude threshold value set in advance according to the statistical analysis of the natural activity amplitude of the body part of the participant in the normal conference scenario, the preset instruction information refers to the corresponding operation instruction information set in advance in the instruction information library for different element parts and their possible change conditions, and the instruction adjustment information refers to the corresponding preset instruction information retrieved from the preset instruction information library according to the element part when the part change amplitude of the element part is greater than the preset part amplitude.

[0117] S3323, when the change state is determined to be combined change, the corresponding element part is regarded as a combined part, and when the combined part has an intersection, the preset instruction information corresponding to the combined part is used as the instruction adjustment information.

[0118] Specifically, when the change state is determined to be combined change, the multiple element parts that change cooperatively can be regarded as combined parts. Then, by analyzing the spatial relationship and action cooperation between these combined parts, it is determined whether the combined parts have an intersection. For example, when a participant squints his eyes while shielding his eyes with his hands, the hands and eyes can be regarded as combined parts. Their changes are related in time and space, and jointly indicate that the light at this moment may be a little dazzling. At this time, the corresponding preset instruction information can be retrieved from the preset instruction information library according to the characteristics of the combined part, such as "please close the curtain", which is used as the instruction adjustment information. Through the analysis of combined change, the complex demand signal of the participant can be more comprehensively and accurately understood, so as to provide a service instruction that is more in line with the actual demand.

[0119] Among them, the combined part refers to the multiple element parts that change their actions, postures or expressions in coordination when the change state of the corresponding element part of the main element (such as the participants) in the video data is judged as a combined change.

[0120] S333, Obtain the display attributes of accompanying elements in the video data, wherein the display attributes include transparency attributes and concealment attributes.

[0121] Specifically, the display attributes of accompanying elements can be determined by analyzing their visual presentation in the video frame. The transparency attribute corresponds to items with transparent materials, such as a glass water cup, whose outline and internal state (such as the water level) can be directly identified through the video. The concealment attribute corresponds to items with opaque materials, such as a stainless steel thermos. Here, display attributes refer to the visual characteristics of accompanying elements in the video data. They can be used to describe the observable and identifiable state of accompanying elements in the video frame. The transparency attribute refers to the display attribute of accompanying elements whose materials are transparent, and whose outline and internal state can be clearly identified and observed directly through the video. The concealment attribute refers to the display attribute of accompanying elements whose materials are opaque, and whose outline and internal state cannot be clearly identified and observed directly through the video.

[0122] S334, when the display attribute is determined to be transparent, the corresponding accompanying element is taken as an identifiable element, the usage height of the identifiable element is identified, and when the usage height is less than the limit height, displacement adjustment information is generated.

[0123] Specifically, when the display attribute of an accompanying element is transparent, it can be identified as a recognizable element. Further quantitative analysis can determine the usage height corresponding to the recognizable element. For example, taking a water cup as an example, by measuring the vertical position of the remaining water in the cup in the video frame, combined with the known frame ratio and actual scene size, its usage height (i.e., the height of the remaining water in the actual scene) can be calculated. Then, a limit height can be preset. When it is determined that the usage height of the recognizable element is less than the limit height, for example, when the height corresponding to the remaining water is significantly lower, it is speculated that the water in the cup is about to run out. At this time, replacement adjustment information can be generated, such as "replenish the beverage for this participant". By monitoring and analyzing the usage height of recognizable accompanying elements, the potential item usage needs of participants can be captured, and corresponding services can be provided in a timely manner.

[0124] Wherein, the identifiable element refers to the accompanying element with the exhibition attribute of transparent attribute, the use height refers to the vertical height of the internal substance (such as water in the water cup) corresponding to the identifiable element in the actual scene, the limit height refers to a critical value preset for measuring whether the use height of the identifiable element reaches the critical value requiring action, and the management adjustment information refers to instruction information generated according to specific circumstances for indicating the operation of replacing or supplementing the corresponding article when the use height of the identifiable element is less than the limit height.

[0125] S335, when the exhibition attribute is the hidden attribute, the corresponding accompanying element is regarded as an unidentifiable element, the use inclination of the unidentifiable element is identified, and when the use inclination is greater than the limit inclination, the adding adjustment information is generated.

[0126] Specifically, if the exhibition attribute of the accompanying element is the hidden attribute, the accompanying element can be marked as an unidentifiable element, the contour and posture of the unidentifiable element can be identified through edge detection, shape fitting and the like, and then the use inclination (i.e. the angle between the article and the horizontal or vertical direction) of the unidentifiable element can be calculated. Similarly, a limit inclination can be preset for measuring whether the inclination state of the article is abnormal. When the use inclination of the unidentifiable element is greater than the limit inclination, such as when the inclination of the opaque water cup is greater than the limit inclination, it can be inferred that the corresponding water cup may be empty, and the corresponding adding adjustment information can be generated to remind the staff to add water for the corresponding participant.

[0127] Wherein, the unidentifiable element refers to the accompanying element with the exhibition attribute of hidden attribute, the use inclination refers to the angle that can be used to reflect the inclination state of the unidentifiable element, the limit inclination refers to a critical value preset for measuring whether the use inclination of the unidentifiable element reaches an abnormal degree, and the adding adjustment information refers to instruction information generated according to the inferred possible situation for indicating the operation of adding the corresponding article when the use inclination of the unidentifiable element is greater than the limit inclination.

[0128] Through the above embodiments, the demand signals of the participants can be comprehensively captured, including the explicit demands expressed through the input data and the potential demands expressed through the output data, so that the accuracy and comprehensiveness of the demand identification of the participants can be improved.

[0129] S4, sending the active adjustment information and the trigger adjustment information to the adjustment device of the target element.

[0130] After obtaining the active adjustment information and the trigger adjustment information, the corresponding adjustment information can be sent to the adjustment device corresponding to the target element, each target element has a corresponding adjustment device, and the adjustment device of the target element can be supervised by the manager of the conference place, when the adjustment device of the target element receives the corresponding adjustment information, the manager can process the received adjustment information, and provide corresponding assistance to the demand of the corresponding target element, for example, when the adjustment information received by the adjustment device of the target element is "replace the drinking water for the person", the manager can provide new drinking water for the corresponding target element. Wherein, the adjustment device refers to a device that can process the analysis obtained adjustment information.

[0131] Referring to Figure 3 , a structure schematic diagram of a mobile terminal edge intelligent inference system supporting multi-modal input is provided, and the data processing system of the mobile terminal edge intelligent inference system based on the multi-modal input supporting includes: The splicing module is configured to receive target images collected by the monitoring device on the target area in response to the pre-completion information, identify the static planning area and the dynamic splicing area of each target element in the target image, and splice to obtain the target identification area of the target element. The calling module is configured to take the corresponding target identification area as a trigger identification area based on the target amplitude of the target element in the target identification area, and call the multi-modal data of the trigger identification area, wherein the multi-modal data includes input data and output data. The analysis module is configured to perform data analysis on the input data to obtain active adjustment information corresponding to the input data, identify the dynamic changes of the subject element and the accompanying element in the output data based on the change identification strategy, and obtain trigger adjustment information of the output data. The sending module is configured to send the active adjustment information and the trigger adjustment information to the adjustment device of the target element.

[0132] Figure 4 The device of the embodiment shown can be used to execute the steps in the method embodiment shown, and the implementation principles and technical effects are similar, which will not be repeated here. Figure 1 The steps in the method embodiment shown, and the implementation principles and technical effects are similar, which will not be repeated here.

[0133] Referring to Figure 4 , a hardware structure schematic diagram of an electronic device is provided, and the electronic device 40 includes a processor 41, a memory 42 and a computer program; wherein The memory 42 is configured to store the computer program, and the memory can also be a flash memory. The computer program is, for example, an application program, a functional module and the like for implementing the above method.

[0134] The processor 41 is configured to execute the computer programs stored in the memory to implement the steps performed by the device in the above method. Details can be referred to the related description in the method embodiments.

[0135] Optionally, the memory 42 can be independent or integrated with the processor 41.

[0136] When the memory 42 is independent of the processor 41, the device can further include: The bus 43 is configured to connect the memory 42 and the processor 41.

[0137] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A mobile edge intelligent inference method supporting multimodal input, characterized in that, include: In response to the pre-completion information, the system receives the target image collected by the monitoring equipment for the target area, identifies and stitches together the static planning area and dynamic stitching area of ​​each target element in the target image to obtain the target recognition area of ​​the target element. Based on the target amplitude of the target element in the target recognition area, the corresponding target recognition area is used as the trigger recognition area, and the multimodal data of the trigger recognition area is retrieved. The multimodal data includes input data and output data. The input data is parsed to obtain active adjustment information corresponding to the input data. Based on the change recognition strategy, the dynamic changes of the main element and the accompanying element in the output data are identified to obtain the trigger adjustment information of the output data. The active adjustment information and the triggered adjustment information are sent to the adjustment device of the target element.

2. The method according to claim 1, characterized in that, The process of identifying and stitching together the static planning region and dynamic stitching region of each target element in the target image to obtain the target recognition region of the target element includes: Identify the positional contours of statically positioned elements and the placement contours of statically placed elements in the target image; Determine the position contour that is closest to each placement side length in the placement contour as the processing contour for each placement side length; The fixed side length and extended side length in the processing contour are processed to obtain the static planning area of ​​the target element; Based on the placement side length, the corresponding static planning area is mirrored to obtain the copied planning area; The copy planning area is dynamically filled in according to the accompanying outline of the accompanying placement element in the placement outline to obtain the dynamic splicing area of ​​the target element. The static planning area and the dynamic splicing area are spliced ​​together to obtain the target recognition area of ​​the target element.

3. The method according to claim 2, characterized in that, The process of processing the fixed side length and extended side length in the processing contour to obtain the static planning area of ​​the target element includes: The side length parallel to the corresponding placement side length in the processing contour is taken as the fixed side length, and the remaining side length is taken as the extended side length. Delete the fixed side length closest to the corresponding placement side length, and extend the extended side length to one side of the corresponding placement side length until it intersects with the placement side length, thus obtaining the static planning area of ​​the target element.

4. The method according to claim 2, characterized in that, The step of dynamically filling the copy planning area according to the accompanying outline of the accompanying placement element in the placement outline to obtain the dynamic splicing area of ​​the target element includes: Identify the accompanying contours of the accompanying placement elements in the placement contour, take the accompanying contours that intersect with the copy planning area as the merged contours, and take the remaining accompanying contours as the contours to be selected. The first combined contour is obtained by the union of the merged contour and the corresponding replicated planning area; The target image is processed by coordinate conversion to obtain the straight-line distance between the center point of the contour to be selected and the first combined contour. The contour to be selected whose straight-line distance is less than the preset filtering distance is selected as the corresponding first combined contour. Obtain the extreme coordinate values ​​of the selected contour and the first combined contour, construct a reconstructed rectangle based on the extreme coordinate values, and use the reconstructed rectangle as the dynamic splicing area of ​​the target element.

5. The method according to claim 1, characterized in that, The step of using the target recognition area as the trigger recognition area based on the target amplitude of the target element in the target recognition area includes: Identify the target amplitude of the target element in the target identification area, and when the target amplitude is greater than a preset amplitude, use the corresponding target identification area as the trigger identification area.

6. The method according to claim 2, characterized in that, The change recognition strategy identifies the dynamic changes of the main elements and accompanying elements in the output data to obtain the triggering adjustment information of the output data, including: Acquire output data, which includes video data and audio data; Based on the dynamic changes of the main elements and accompanying elements in the audio data using the audio recognition strategy, the triggering and adjustment information of the audio data is obtained. Based on the video recognition strategy, the dynamic changes of the main elements and accompanying elements in the video data are analyzed to obtain the trigger adjustment information of the video data.

7. The method according to claim 6, characterized in that, The method based on audio recognition strategy to dynamically change the main elements and accompanying elements in audio data, and to obtain the trigger adjustment information of audio data, includes: Based on timbre recognition, the main sound wave in the audio data is taken as the main element, and the remaining sound waves are taken as accompanying elements; Extract the main amplitude of the main element in the audio data, and when the main amplitude is greater than the preset main amplitude, take the corresponding sound wave as the main change sound wave, and obtain the main duration of each main change sound wave; When the duration of the main body is determined to be less than the preset duration, the corresponding change in the main body sound wave is taken as the instantaneous main body sound wave, and the preset transmission information is retrieved based on the instantaneous main body sound wave; When the duration of the main body is determined to be greater than the preset duration, the corresponding change sound wave of the main body is taken as the continuous main body sound wave, the text information corresponding to the continuous main body sound wave is extracted, and annotation adjustment information is generated based on the text information; Extract the accompanying amplitude of the accompanying element in the audio data. When the accompanying amplitude is greater than the preset accompanying amplitude, the corresponding sound wave is taken as the accompanying change sound wave, and the accompanying duration of each accompanying change sound wave is obtained. When it is determined that the accompanying duration is less than the preset duration, the corresponding accompanying change sound wave is taken as the instantaneous accompanying sound wave, and the processing and adjustment information is retrieved based on the instantaneous accompanying sound wave. When the duration of the accompanying sound wave is determined to be longer than a preset duration, the corresponding accompanying sound wave is taken as a continuous accompanying sound wave, and a cancellation sound wave opposite to the continuous accompanying sound wave is generated. The cancellation sound wave is used as cancellation adjustment information.

8. The method according to claim 6, characterized in that, The step of obtaining trigger adjustment information for video data based on the dynamic changes of main elements and accompanying elements in the video data according to the video recognition strategy includes: Identify target elements in video data as main elements, and accompanying elements as accompanying elements; Identify the changing state of the corresponding element parts of the main element in the video data, and generate instruction adjustment information based on the changing state; Obtain the display attributes of accompanying elements in the video data, including transparency and concealment attributes; When the display attribute is determined to be transparent, the corresponding accompanying element is taken as an identifiable element, the usage height of the identifiable element is identified, and when the usage height is less than the limit height, displacement adjustment information is generated. When the display attribute is determined to be a hidden attribute, the corresponding accompanying element is treated as an unrecognizable element. The usage tilt of the unrecognizable element is identified. If the usage tilt is greater than the limit tilt, adjustment information is generated.

9. The method according to claim 8, characterized in that, The method of identifying the changing state of the corresponding element parts of the main element in the video data, and generating instruction adjustment information based on the changing state, includes: Identify the changing state of the corresponding element parts of the main element in the video data, wherein the changing state includes independent changes and combined changes; When the change state is determined to be an independent change, the change range of the element part is identified. When the change range of the part is determined to be greater than the preset part range, the corresponding preset instruction information is retrieved based on the element part as instruction adjustment information. When the change state is determined to be a combination change, the corresponding element part is taken as the combination part. When the combination parts have an intersection, the preset instruction information corresponding to the corresponding combination part is taken as the instruction adjustment information.

10. A mobile edge intelligent inference system supporting multimodal input, characterized in that, include: The stitching module is used to respond to the pre-completion information, receive the target image collected by the monitoring device on the target area, identify the static planning area and dynamic stitching area of ​​each target element in the target image and stitch them together to obtain the target recognition area of ​​the target element; The retrieval module is used to retrieve multimodal data of the trigger recognition area based on the target amplitude of the target element in the target recognition area, and the corresponding target recognition area is used as the trigger recognition area. The multimodal data includes input data and output data. The parsing module is used to parse the input data to obtain active adjustment information corresponding to the input data, and to identify the dynamic changes of the main elements and accompanying elements in the output data based on the change recognition strategy to obtain the trigger adjustment information of the output data. The sending module is used to send the active adjustment information and the triggered adjustment information to the adjustment device of the target element.

Citation Information

Patent Citations

  • System and method for processing visual, auditory, olfactory, and / or haptic information

    CN103003761A

  • Conference scene picture tracking method and device and storage medium

    CN116866509A

  • Data processing method based on large model

    CN119623774A

  • Video conferencing system with physical cues

    US20050110867A1