A three-port semantic communication architecture optimization method introducing a semantic mapping terminal

By introducing a semantic mapping endpoint into the semantic communication system, semantic vectors are monitored and mapped in real time, solving the problem of semantic uninterpretability in existing systems. This enables efficient and reliable semantic information transmission and processing, meeting the needs of complex application scenarios.

CN119675827BActive Publication Date: 2025-12-12NANTONG RES INST FOR ADVANCED COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411724352.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-12-12
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Most existing semantic communication systems adopt a two-end structure, which ignores the monitoring and interpretation of semantic information, leading to problems such as uninterpretable or inconsistent semantics. This is especially true in the fields of artificial intelligence and automated decision-making, affecting communication quality and reliability.

Method used

By introducing a semantic mapping end, a multimodal semantic encoding and decoding model is constructed to monitor and map semantic vectors in real time, providing interpretability and transparency, identifying semantic loss during transmission and correcting semantic distortion, thereby increasing the reliability and controllability of semantic communication.

Benefits of technology

It improves the robustness and reliability of semantic communication, ensuring that the core semantics of information are preserved and conveyed to the maximum extent under imperfect network conditions, adapting to the needs of complex application scenarios, and enhancing the overall performance and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119675827B_ABST
    Figure CN119675827B_ABST
Patent Text Reader

Abstract

The application provides a three-port semantic communication architecture optimization method with a semantic mapping end, and belongs to the technical field of information processing. The technical problems of semantic communication service content being not conducive to explanation, supervision and integration are solved. The technical scheme comprises the following steps: S1: constructing and training a multi-modal semantic encoding model; S2: constructing and training a multi-modal semantic decoding model; S3: constructing a fixed scene, and testing the precision and dimension of a high compression floating point format vector; S4: fixing a service scene; S5: collecting service original multi-modal data; S6: collecting corresponding encoding data of the service original multi-modal data according to step S4; S7: constructing and training a mapping model of the encoding data to text or label data; S8: realizing a normal semantic communication system transmission process; and S9: performing subsequent supervision or processing operations. More research and optimization space is provided for future semantic communication systems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, and in particular to a three-port semantic communication architecture optimization method introducing a semantic mapping end. BACKGROUND

[0002] With the continuous development of communication technology, semantic communication has gradually become an important research direction in the field of information transmission. Traditional communication networks usually focus on accurate transmission of signals, i.e. bit-level information transmission. However, with the explosive growth of information volume and the widespread popularity of intelligent applications, merely transmitting bit information cannot meet the needs in complex application scenarios, especially in cases where information content needs to be understood and interpreted. Semantic communication significantly improves communication efficiency and content understanding ability by converting raw information into higher-level semantic representation, and becomes an important technology development trend for next-generation communication networks.

[0003] However, existing semantic communication systems mostly adopt a two-end structure, i.e. one sending end and one receiving end. The sending end encodes information into semantic vectors, and correspondingly, the receiving end decodes and reconstructs the original content. This design structure to some extent ignores the monitoring and interpretation of semantic information in the communication process, leading to undesirable results such as uninterpretable or inconsistent semantics in actual application, especially in the field involving artificial intelligence and automated decision-making.

[0004] In semantic communication systems, semantic loss or semantic distortion is an important problem affecting communication quality and system reliability. Semantic loss refers to the failure to successfully transmit or ignore part of the semantic information in the information transmission process, while semantic distortion refers to the misinterpretation of transmitted semantic information at the receiving end or deviation from the original intention. These problems can lead to the inability of the information receiver to accurately understand the sender's intention, especially in applications involving automated decision-making, intelligent control or human-computer interaction, semantic distortion can have serious consequences.

[0005] Therefore, the current semantic communication technology needs a new implementation to ensure efficient transmission while controlling and processing semantic information. SUMMARY

[0006] The purpose of the present application is to provide a three-port semantic communication architecture optimization method introducing a semantic mapping end, which solves the two problems above, i.e. the semantic channel state is not transparent in the semantic communication of the prior art, and cannot meet the needs of semantic control and processing in complex scenarios; in addition, the semantic vector bandwidth occupation in the semantic channel transmission of the prior art is difficult to achieve optimization, and there is a certain optimization space in the scenario considering only transmission as the purpose, therefore, the purpose of the present application is to solve the above two problems and realize the optimization of the semantic communication architecture.

[0007] The invention idea of the present application is: starting from explainability, the problem of semantic loss and semantic distortion is effectively solved by introducing a semantic mapping end. First, the semantic mapping end monitors and maps the semantic vector in real time while transmitting it, converting the floating-point array into interpretable semantic information. This mapping not only provides an intuitive understanding of the transmitted content, but also allows for timely detection and correction of possible semantic distortion by detecting the difference between the mapping result and the expected semantics. This three-port semantic communication network with enhanced explainability adds a semantic mapping end to the traditional transceiver, which performs interpretable semantic mapping on the semantic vector during transmission. By introducing the semantic mapping end, the system can monitor and manage the transmitted semantic information in real time, providing higher explainability and transparency to better adapt to the needs of complex application scenarios. This innovation not only improves the reliability and controllability of semantic communication, but also lays a solid foundation for the development of future intelligent communication systems. Through monitoring and management of semantic information, the semantic mapping end can identify semantic loss during transmission and notify the sender to re-encode or supplement the transmission through a feedback mechanism. This mechanism greatly improves the robustness of the communication system, ensuring that even under less-than-ideal network conditions, the core semantics of the information can be preserved and conveyed to the greatest extent possible.

[0008] To achieve the above-mentioned application purpose, the technical scheme adopted by the present application is specifically as follows: a three-port semantic communication architecture optimization method introducing a semantic mapping end, comprising the following steps:

[0009] S1: Construct and train a multi-modal semantic encoding model to convert the multi-modal information transmitted by the channel into a high-compression floating-point format vector that is conducive to transmission. Here, compression refers to the conversion of multi-modal information into a one-dimensional vector form with a limited value range that is easy to process by a computer through flattening, normalization, standardization, and other operations. Then, the vector is forward propagated through a neural network model, and according to different scenarios and data volumes, it is output as a high-compression floating-point format vector with a relatively small data volume, containing the amount of information of interest to the business and the amount of information required for multi-modal recovery;

[0010] S2: Construct and train a multi-modal semantic decoding model to restore the floating-point format vector in step S1 to the original multi-modal data form as much as possible without loss;

[0011] Specifically, the high-compression floating-point format vector obtained in step S1 will continue to be flattened as input to the subsequent neural network, and forward propagation through the neural network model will result in a vector with increased data volume that directly corresponds to the original input multi-modal information. Then, according to the composition of the vector, the data is re-integrated, classified, and other operations to restore it to multi-modal data forms such as voice, video, and text;

[0012] S3: Constructing a fixed scene, the accuracy and dimensionality requirements of the high-compression floating-point format vector are tested, and the highest efficient data accuracy and dimensionality are found under the condition that the multi-modal codec restores the data effect to adapt to the business demand;

[0013] For example, for occasions that limit the compression rate, if the required data compression rate is 50%, the vector dimension of the output part in step S1 can be set to 50% of the input layer data amount. This index can be achieved by data quantization or reduction of the number of neurons, and on this basis, the restoration effect of step S2 output is optimized, that is, the transmission quality is optimized as much as possible under the condition of limiting bandwidth.

[0014] If it is for a scene that requires restoration effect, for example, the restoration rate is 80%, the multi-modal data restored by step S2 output is compared with the original data and quantified by cross-entropy loss. On the basis of a large amount of training, ensure that the restoration rate can finally stabilize at more than 80%, and try to limit the amount of intermediate data through data quantization or reduction of the number of neurons, so as to find the most efficient data accuracy and dimensionality. The quantification of the restored data effect can follow the following formula:

[0015]

[0016] Where: X is the original input data, Y is the data restored by the model, and R is the restoration rate based on relative error;

[0017] S4: Fix the business scene and conduct a certain amount of business testing to collect business data; Here the data refers to the multi-modal data input in step S1. If it is a monitoring type of business scene, the monitoring video needs to be collected as business data;

[0018] S5: According to the collected business original multi-modal data, the industry experts label the text or label of the multi-modal business original data; For example, when an event or phenomenon of business concern occurs in the monitoring, the expert needs to mark and label the corresponding data part;

[0019] S6: According to the encoding data corresponding to the business original multi-modal data collected in step S4, construct a binary data set with the text or label in step S5, so as to complete the mapping of multi-modal data to business information, which is convenient for subsequent sorting and training;

[0020] S7: For the binary data set collected under the fixed business scenario, a mapping model of encoding data to text or label data is constructed and trained, and the training is continuously carried out as the business data increases; here, according to the data format, data size, label quantity and other information, a neural network with corresponding complexity and a customized training strategy are designed by an artificial intelligence expert to ensure the orderly training;

[0021] S8: Through the model trained in step S7, additional semantic ports are derived and semantic text or label data is output in the transmission process of the normal semantic communication system; here, the derived semantic port can be considered as an output port directly derived from a hidden layer of the overall neural network, and both the derived semantic port and the multi-modal restoration information output in step S2 are the optimization targets, but the semantic port and the multi-modal restoration information are two independent ports and are not directly related;

[0022] The optimization of the multi-modal restoration information can follow the loss function based on the following rules:

[0023]

[0024] L compress =D KL (q(Z|X)||p(Z))

[0025] L=L recon +β·L compress

[0026] Wherein, L recon is the reconstruction error loss function, X is the original input data, Y is the data restored by the model, and P is the norm type. The loss function measures the difference between the restored data and the original input data, and the goal is to minimize the reconstruction error as much as possible.

[0027] Wherein L compress is the semantic bandwidth loss function, p(Z) is the prior distribution in the form of standard normal distribution, X is the original input data, q(Z|X) is the posterior probability distribution of the latent representation based on the input X, D KL is the Kullback-Leibler divergence of the prior and posterior distributions, i.e. the distance between the distributions. This term is used to constrain the compactness of the latent semantic representation, i.e. to limit the channel bandwidth.

[0028] Wherein L is the overall loss function, β is the balance coefficient, used to balance the channel bandwidth (compression degree) and the restoration quality. By optimizing the loss function with the goal of minimizing L, a good balance between channel width and restoration quality is expected.

[0029] S9: Collecting the semantic text or label derived in step S8 by the control system, and performing subsequent supervision or processing operations. When the semantic communication system in the production environment, the alarm type label appears in the semantic port, the staff can flexibly process according to the label content, for example, the unimportant alarm type can be ignored, and the urgent alarm type can directly apply to view the recovered multi-modal information, so as to flexibly perform subsequent operations, avoid frequent processing of a large amount of multi-modal information, and cause personnel fatigue or hardware resource occupation.

[0030] Compared with the prior art, the beneficial effects of the present application are:

[0031] 1. The actual content in the transmission channel of the semantic communication is related to the semantic of the multi-modal content, but not identical, the actual content in the channel should be smaller in data quantity than the multi-modal content, thereby reflecting the role of semantic communication in bandwidth saving, if the compression rate of semantic communication is 10% under the same bandwidth network environment, the amount of multi-modal information that can be transmitted is increased to 10 times of the original, that is, more multi-modal information can be transmitted under the same condition, or the cost of bandwidth is correspondingly reduced under the same business volume.

[0032] 2. The actual content in the transmission channel of the present application not only pays attention to the bandwidth of transmission, but also considers the integrity of the key semantic features, since the multi-modal information recovered by the semantic communication needs to match the original input multi-modal information as much as possible, therefore the actual content in the transmission channel must contain necessary semantic information, such as change of object state, occurrence of event, scene switching, etc. Once the bandwidth of the actual content in the transmission channel is excessively limited for compression rate, part of the important semantic information may be lost, resulting in unacceptable distortion of the recovered multi-modal information.

[0033] 3. The redundant semantic features of the present application should be optimized by the encoder and decoder and not participate in transmission. For example, the object properties that remain unchanged for a long time can not participate in transmission, and the secondary features irrelevant to the business focus can also not participate in transmission, thereby more conducive to maintaining a small transmission bandwidth.

[0034] 4. The semantic text and label derived by the third port of the present application should be in a general format that is standard, easy for third parties to directly understand and process, formed on the basis of the actual content of transmission and integrated with redundant semantic features, for example, for the movement of objects in video business, relatively general and easy-to-understand label formats for operators can be used, such as up, down, left, right, etc. If necessary, mixed descriptions of multiple types should not be used, such as lifting, pulling up, lowering, horizontal height increase, etc. The use of standard expression labels can improve the reaction speed of operators, or facilitate the archiving, integration, statistics, etc. of semantic labels.

[0035] 5、The semantic text and tags of the third port of the application can be used for the automatic control process of the integrated business scenario, and subsequent supervision or processing operations can be performed. The semantic text and tags can be exported in a relatively standardized json format, including objects, features, attributes, etc. containing semantic information, so as to avoid the need for re-wording, classification or more complex natural language processing in subsequent processing.

[0036] 6、The important advantage of the application lies in its flexibility and scalability. By introducing the semantic mapping end, the system can not only adapt to different application scenarios, but also be customized according to specific needs. For example, in a security-sensitive communication environment, the semantic mapping end can be configured to detect potential malicious content or abnormal semantics in real time, thereby providing additional security protection. In cross-language communication scenarios, the semantic mapping end can act as an intermediate layer to realize real-time translation at the semantic level, further expanding the application range of semantic communication.

[0037] 7、The three-port architecture of the application provides more research and optimization space for future semantic communication systems. For example, machine learning algorithms can be integrated into the semantic mapping end to enable adaptive optimization of the semantic mapping process, improving the overall performance and adaptability of the system. This design not only improves the interpretability of the system, but also provides convenience for subsequent performance optimization and function expansion. BRIEF DESCRIPTION OF DRAWINGS

[0038] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, which together with the embodiments of the application, are used to explain the application, and do not constitute a limitation on the application.

[0039] Figure 1 The network structure of the three-port semantic communication architecture of the application from the input of business data to the semantic port and the output port. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical scheme and advantages of the application more clear and explicit, the application will be further described in detail below in combination with the drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the application, and do not limit the application.

[0041] Embodiment 1

[0042] Semantic communication example of pipeline video monitoring system

[0043] System overview

[0044] In modern intelligent manufacturing environments, real-time monitoring of production lines is crucial for ensuring product quality and production efficiency. The three-port semantic communication network proposed in this invention can be effectively applied to the flow line video monitoring system, achieving efficient and low-bandwidth real-time monitoring while providing intelligent early warning functions.

[0045] This embodiment optimizes the video monitoring system on the manufacturing flow line. Traditional methods usually directly transmit high-definition video streams, which not only occupies a large amount of bandwidth, but also makes it difficult to achieve real-time intelligent analysis. This system uses the principle of semantic communication to convert video information into abstract semantic vectors directly related to product status, greatly reducing data transmission volume while improving the system's intelligent analysis capabilities.

[0046] Sender design

[0047] 1. Data acquisition module

[0048] Deploy high-resolution cameras to cover key positions on the flow line.

[0049] The frame rate can be adjusted (e.g., 30fps) to adapt to different monitoring needs.

[0050] 2. Preprocessing unit

[0051] Preprocess the original video stream, such as denoising and contrast enhancement.

[0052] Use computer vision algorithms to extract key information from each frame, such as product contours, position coordinates, and motion vectors.

[0053] 3. Semantic encoder

[0054] Design a special deep neural network to encode the preprocessed video frame information into a fixed-dimensional semantic vector. The semantic vector is related to product posture, speed, position, and other key state information, and is not original pixel data or key information itself, but a high-dimensional compressed representation of key information.

[0055] For example, a floating-point data that maps to the sum of the mean square error loss of multiple product trajectories on the flow line and the preset trajectory does not represent the running trajectory of a specific object, but contains constraints on the motion characteristics of each object.

[0056] The floating-point data obtained by actually training the neural network may be more abstract high-dimensional features that are positively correlated with a large number of physical features, but do not correspond to specific physical meanings.

[0057] Compression of semantic vectors can use quantization techniques or efficient encoding algorithms.

[0058] Compression rate can be dynamically adjusted to balance transmission efficiency and information fidelity.

[0059] Receiver design

[0060] 1. Data decompression module

[0061] Receive compressed semantic vectors and decompress them.

[0062] Implement automatic recognition and decoding of various compression algorithms.

[0063] 2. Semantic decoder

[0064] Use a deep neural network structure that matches the encoder at the sending end.

[0065] Reconstruct the decompressed semantic vectors into a sequence of video frames, realizing the mapping of semantic encoding and pipeline monitoring video.

[0066] 3. Video reconstruction module

[0067] Based on the reconstructed frame sequence, generate a continuous video stream.

[0068] Apply interpolation algorithms to improve the smoothness and realism of the reconstructed video.

[0069] 4. Display output unit

[0070] Support multiple display devices such as monitoring screens, mobile terminals, etc.

[0071] Provide a user-friendly interface that supports multi-angle switching, picture enlargement, etc.

[0072] Semantic mapping end design

[0073] 1. Vector capture module

[0074] Capture semantic vectors in real time during transmission.

[0075] Implement compatibility with different network protocols to ensure data integrity.

[0076] 2. Semantic mapping model

[0077] Industry experts label multi-dimensional information such as product posture, position deviation, speed anomaly, etc. in business video, and form binary tuples with video frames as training materials.

[0078] Build a special neural network model to map abstract semantic vectors to specific product state descriptions.

[0079] The model output includes semantic text or semantic labels of multi-dimensional information such as product posture, position deviation, speed anomaly, etc.

[0080] 3. State analysis unit

[0081] Based on the mapping results, real-time state assessment is conducted. The assessment can be obtained by directly inputting the semantic text or labels into a large language model familiar with industry processes, resulting in an assessment of the current state, such as whether there are abnormalities, whether intervention is needed, how to handle it, etc. Multiple threshold levels can be set to distinguish between normal state, slight abnormality, and serious failure.

[0082] 4. Alarm generation module

[0083] When an abnormal state is detected, an alarm information of corresponding level is automatically generated.

[0084] Supports multiple alarm methods, such as system pop-up, SMS notification, sound and light alarm, etc.

[0085] The alarm system can cascade corresponding emergency handling facilities, such as triggering emergency stop button, power-off button, fire switch, etc. according to actual emergency situation.

[0086] System workflow

[0087] 1. Normal monitoring process

[0088] (1) The sending end camera continuously captures pipeline video.

[0089] (2) The pre-processing unit extracts key information, and the semantic encoder generates semantic vectors.

[0090] (3) The compressed vectors are transmitted to the receiving end and the semantic mapping end through the network.

[0091] (4) The receiving end reconstructs the video picture for real-time monitoring by the operator.

[0092] (5) The semantic mapping end converts the vectors into state descriptions, such as "product posture normal, conveyor running smoothly".

[0093] 2. Abnormality detection process

[0094] (1) The semantic mapping end detects an abnormality, such as "product A deviates from the center line by 2.5 cm".

[0095] (2) The state analysis unit assesses the degree of abnormality and determines that intervention is needed.

[0096] (3) The alarm generation module immediately sends an alarm to the relevant personnel.

[0097] (4) At the same time, the system can automatically adjust the display of the receiving end and highlight the abnormal area.

[0098] System advantages

[0099] 1. High efficiency transmission

[0100] Compared with traditional video streaming, semantic vector transmission can reduce bandwidth occupancy by more than 90%. Support more cameras online at the same time, improve monitoring coverage.

[0101] 2、Intelligent analysis capability

[0102] The semantic mapping end realizes intelligent interpretation from abstract data to specific state.

[0103] Support anomaly detection in complex scenarios, such as multiple products deviating from the predetermined trajectory at the same time.

[0104] 3、Real-time response

[0105] Low-latency data transmission and processing ensure that abnormal conditions can be discovered and handled in a timely manner. Support millisecond-level alarm triggering to minimize production loss.

[0106] 4、System scalability

[0107] Semantic encoding and mapping models can be continuously optimized through continuous learning.

[0108] Easy to integrate other sensor data such as temperature, vibration, etc., to realize multi-modal monitoring. Compared with traditional methods, as shown in Table 1

[0109] Table 1

[0110]

[0111]

[0112] This embodiment demonstrates the innovative application of a three-port semantic communication network in the field of intelligent manufacturing. By compressing high-dimensional video information into low-dimensional semantic vectors, not only is the transmission efficiency greatly improved, but also an ideal data foundation is provided for intelligent analysis and early warning. The introduction of the semantic mapping end enables the system to automatically interpret complex product states, achieving a leap from passive monitoring to active early warning. This method provides a new approach for the intelligent upgrading of manufacturing, and is expected to play an important role in improving production efficiency and reducing human intervention. In practical application scenarios such as intelligent factories and cross-language communication, this invention has shown significant advantages, paving the way for future more intelligent and safer communication systems.

[0113] Embodiment 2

[0114] Farm acoustic monitoring system embodiment

[0115] System overview

[0116] This embodiment is aimed at the optimization design of the acoustic monitoring system of modern farms. The traditional method usually uses continuous recording or sampling recording, which not only produces a large amount of data that needs to be stored, but also makes it difficult to realize real-time analysis and early warning of poultry sounds. This system uses the principle of semantic communication to convert audio information into descriptive semantic tags, which not only significantly reduces data transmission volume, but also enables intelligent identification and analysis of poultry calls and device sounds.

[0117] Sender design

[0118] 1. Data acquisition module

[0119] Deploy omnidirectional microphone arrays at key locations in the farm;

[0120] Adjustable sampling rate (e.g. 44.1 kHz) to ensure collection of sound details;

[0121] Supports sound source positioning function to determine the location of the sound source.

[0122] 2. Preprocessing unit

[0123] Perform ambient noise elimination and sound enhancement processing

[0124] Use spectrogram analysis technology to extract audio features

[0125] Perform sound framing and feature extraction, such as Mel Frequency Cepstral Coefficients (MFCC)

[0126] 3. Semantic encoder

[0127] Use a specially trained deep neural network to encode audio features into semantic vectors

[0128] Semantic vectors contain feature information of chicken call types (such as eating, resting, fighting, mating, etc.) and device operating sounds (such as fans, feeders, water supply systems, etc.)

[0129] Use efficient compression algorithms to compress semantic vectors

[0130] Support real-time encoding to ensure the timeliness of the monitoring system.

[0131] Receiver design

[0132] 1. Data decompression module

[0133] Receive and decompress semantic vectors

[0134] Ensure data integrity and real-time performance

[0135] 2. Semantic decoder

[0136] Use a neural network structure that matches the encoder

[0137] Reconstructing semantic vectors into audio signals

[0138] Supporting selective reconstruction, only reconstructing audio of specific time period

[0139] 3. Audio reconstruction module

[0140] Generating continuous audio stream

[0141] Applying audio enhancement algorithm to improve the quality of reconstructed audio

[0142] Supporting multi-channel output

[0143] 4. Playback control unit

[0144] Providing audio playback interface, supporting progress bar dragging, volume adjustment and other functions

[0145] Supporting quick positioning to abnormal audio segment by timestamp

[0146] Providing visual display of audio data (such as spectrogram)

[0147] Semantic mapping end design

[0148] 1. Vector analysis module

[0149] Real-time capture of semantic vectors for audio

[0150] Vector feature extraction and classification

[0151] 2. Semantic label generator

[0152] Mapping semantic vectors to specific sound type descriptions

[0153] Supporting multi-level label system:

[0154] Chicken call classification: normal call, fighting sound, pain sound, courtship sound, etc. Equipment sound classification: normal operation, abnormal vibration, bearing wear, etc. In addition, support generating timestamp and label association record 3. Status evaluation unit

[0155] Statistical analysis based on semantic labels to evaluate group health status (such as fighting frequency, abnormal call proportion) and monitor equipment operating status (such as fan normal operation rate)

[0156] Generating trend report and health index

[0157] 4. Early warning system

[0158] Setting multi-level warning threshold

[0159] When detecting abnormal sound pattern, timely alarm, supporting multiple warning methods such as SMS, application push, etc.

[0160] Recording early warning events and generating analysis reports

[0161] System workflow

[0162] 1. Daily monitoring process

[0163] (1) The microphone array continuously collects audio from the farm

[0164] (2) Generate semantic vectors after preprocessing and transmit

[0165] (3) The receiving end reconstructs the audio on demand for manual review

[0166] (4) The semantic mapping end generates labels and performs analysis

[0167] (5) Generate operation reports regularly

[0168] 2. Abnormal processing process

[0169] (1) The system detects abnormal sound (such as intensive fighting sound)

[0170] (2) Automatically generate abnormal event records, including time and location information (3) Send early warning information to management personnel

[0171] (4) Support quick retrieval of original audio for corresponding period

[0172] (5) Record processing results for system optimization

[0173] System feature functions

[0174] 1. Intelligent analysis function

[0175] Analysis of bird behavior patterns

[0176] Device health status diagnosis

[0177] Environmental anomaly detection

[0178] 2. Statistical analysis function

[0179] Generate daily / weekly / monthly operation reports

[0180] Device operation status statistics

[0181] Analysis of bird health trends

[0182] 3. Retrieval function

[0183] Retrieve sound events by time and type

[0184] Support similar event association query

[0185] Abnormal event trace analysis

[0186] System advantages

[0187] 1. High efficiency operation

[0188] Significant reduction in data storage requirements

[0189] Support for large-scale farm monitoring

[0190] Reduced frequency of manual inspections

[0191] 2. Intelligent early warning

[0192] Timely detection of abnormal situations

[0193] Reduced risk of disease transmission

[0194] Prevention of equipment failure

[0195] 3. Data value

[0196] Support for long-term data analysis

[0197] Assistance in optimizing breeding decisions

[0198] Provide scientific basis for breeding

[0199] Comparison with traditional methods, as shown in Table 2:

[0200] Table 2

[0201]

[0202]

[0203] This embodiment demonstrates the innovative application of a three-port semantic communication network in the field of intelligent breeding. By converting complex acoustic information into structured semantic tags, not only is efficient data transmission and storage achieved, but also a powerful tool for intelligent management of the farm is provided. The semantic mapping function of the system enables farm managers to quickly grasp the situation on the farm, enabling intelligent monitoring of the health of the flock and the operation of equipment, providing a new technical solution for modern breeding.

[0204] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for optimizing a three-port semantic communication architecture with a semantic mapping end introduced, characterized in that, Comprise the following steps: S1: build and train a multi-modal semantic encoding model for translating multi-modal information transmitted by channels into a high-compression floating-point format vector conducive to transmission; S2: build and train a multi-modal semantic decoding model for losslessly restoring the high-compression floating-point format vector in step S1 to the original multi-modal data form; S3: In a fixed scenario, test the accuracy and dimensionality of the high-compression floating-point format vector, and find the most efficient data precision and dimensionality in terms of bandwidth resource occupation when the multi-modal encoder-decoder restores data to meet business demand. The quantification of the restored data effect can follow the following formula: Where: X is the original input data, Y is the data restored by the model, and R is the restoration rate based on relative error; S4: Fix the business scenario and conduct a certain amount of business testing to collect business original multi-modal data; S5: According to the collected business original multi-modal data, the industry experts annotate the multi-modal business original data with text or labels; S6: According to the encoding data corresponding to the business original multi-modal data collected in step S4, construct a binary tuple data set with the text or label annotation in step S5, thereby completing the mapping of multi-modal data to business information; S7: For the binary tuple data set collected in the fixed business scenario, build and train a mapping model from the encoding data to the text or label data, and continuously train as the business data increases; S8: Through the model trained in step S7, additional semantic ports are exported and semantic text or label data is output during the transmission process of the normal semantic communication system; S9: The semantic text or label exported in step S8 is collected by the control system to perform subsequent supervision or processing operations.

2. The method of claim 1, wherein the three-port semantic communication architecture optimization method with introduced semantic mapping end is characterized in that, The compression in step S1 refers to the flattening, normalization, and standardization of multi-modal information to convert the data into a one-dimensional, value range within a limited range, and easy-to-computer-process vector form. Then, the vector is forward propagated through a neural network model. Depending on the specific scenario and data volume, it is output as a high-compression floating-point format vector with a relatively small data volume, containing the amount of information of interest to the business and the amount of information needed to restore the multi-modal.

3. The method of claim 1, wherein the three-port semantic communication architecture optimization method with introduced semantic mapping end is characterized in that, In step S2, the high-compression floating-point format vector obtained in step S1 is continued in a flattened form as input to the subsequent neural network, and is forward propagated through a neural network model to obtain a vector with increased data volume that directly corresponds to the original input multi-modal information. Then, according to the composition of the vector, the data is re-integrated and classified to restore the multi-modal data form of voice, video, and text.

4. The method of claim 1, wherein the three-port semantic communication architecture optimization method with introduced semantic mapping end is characterized in that, In step S3, the vector dimension of the output part in step S1 is set to 50% of the input layer data volume. This target is achieved through data quantization or reduction of the number of neurons, and the restoration effect in step S2 is optimized based on this.

5. The method of claim 1, wherein, The data in the step S4 refers to the multi-modal information input in the step S1, and if it is a monitoring type of business scenario, a large amount of monitoring videos need to be collected as business data.

6. The method of claim 1, wherein, The semantic port derived in the step S8 is an output port directly derived from a hidden layer of the whole neural network, and both the multi-modal recovery information output in the step S2 and the multi-modal recovery information are targets for training and optimization. The optimization of the multi-modal recovery information follows a loss function based on the following rules: L compress = D KL (q(Z|X)||p(Z)) L = L recon + β · L compress wherein L recon is a reconstruction error loss function, X is the original input data, Y is the data recovered by the model, P is a norm type, the loss function measures the difference between the recovered data and the original input data, and the goal is to reduce the reconstruction error; where L compress is the semantic bandwidth loss function, p(Z) is the prior distribution shaped like a standard normal distribution, X is the original input data, q(Z|X) is the posterior probability distribution of the latent representation obtained based on the input X, D KL is the Kullback-Leibler divergence of the prior and posterior distributions, i.e., the distance between the distributions, which is used to constrain the compactness of the latent semantic representation, i.e., to limit the channel bandwidth; Wherein L is the whole loss function, and β is a balance coefficient for balancing the channel bandwidth and the recovery quality, and the minimization of L is taken as the goal, and through the optimization of the loss function, a good balance between the channel width and the recovery quality is expected to be achieved.

7. The method of claim 1, wherein the three-port semantic communication architecture optimization method with introduced semantic mapping end comprises: In the step S9, when the semantic communication system in the production environment, the semantic port appears an alarm type of label, a staff member flexibly handles according to the label content, ignores an unimportant alarm type, directly applies to view the recovered multi-modal information for an urgent alarm type, and thus flexibly executes a subsequent operation.

Citation Information

Patent Citations

  • Cross-modal data retrieval system based on text semantic mapping and retrieval method thereof

    CN110990597A

  • Cross-modal hash retrieval method based on controlled semantic embedding

    CN112948601A