{
    "created": "2025-11-05 19:04:57",
    "updated": "2026-08-07 23:59:05",
    "id": "d840d65e-d096-41cc-ba76-32e304a1b2b8",
    "version": 2,
    "ds_topic": null,
    "title_cn": "新疆10m水体透明度数据集（2024年4-10月）",
    "title_en": "",
    "ds_abstract": "<p>本数据采用2024年哨兵2号光学卫星数据为数据源，采样随机森林回归（RFR）方法构建新疆湖泊水体透明度遥感估算模型。数据为CGCS2000坐标系阿伯斯投影，精度为10米。数值系数为0.001，即数据像素值乘以0.001可得实际水体透明度（m）。</p>",
    "ds_source": "<p>采用2024年哨兵2号光学卫星数据L1C级别数据，数据可从欧空局哥白尼数据开放中心获取（https://dataspace.copernicus.eu/browser/）。</p>",
    "ds_process_way": "<p> 使用具有较好辐射性能的Sentinel 2 A/B MSI。Sentinel 2 A/B MSI拥有13个光谱波段，空间分辨率分别是10 m、20 m和60 m，重访周期为10天，双星组网后为5天，在内陆水体表现了较好的性能。从欧空局哥白尼数据开放中心获取L1C数据。\n采用POLYMER算法进行大气校正。随后需要进一步去除天空光、太阳耀斑和残余气溶胶散射的影响：\nR_rs (λ)=(R(λ)-min⁡(R(865),R(2202)))/π\n式中，Rrs为水体遥感反射率，R为POLYMER校正获取的地表反射率。\n采样随机森林回归（RFR）方法构建水质参数模型。随机森林是一种基本单元为决策树的集成学习方法，在训练过程中构建大量相互独立的决策树形成“随机森林”，最后综合这些决策树结果提高模型精度（例如输出平均值）。RFR算法的“随机性”主要体现在两个方面：构建每颗决策树时通过Bagging方法（即自助采样法，每次采样后将样本放回）从原始训练数据集中随机抽样生成训练数据子集；在节点分裂时不使用所有的特征变量参与比较，而是从特征变量中随机选择一个子集参与节点分裂。两个“随机性”使得RFR算法不容易过拟合，并且对异常值和噪声具有很好的容忍度。\nRFR算法在训练过程中最重要的几个超参数是：（1）决策树的数量（n_estimators）。n_estimators越大，模型效果通常越好，但计算时间也越长；在达到一定数据量，模型趋于稳定。（2）节点分裂时的最大特征数量（max_features），决策树在节点分裂时从随机选择的max_features个特征中寻找最佳分裂特征。（3）决策树的最大深度（max_depth），如果不设置（即None），则决策树会最大限度的生长直到满足分割终止条件。RFR通过Python scikit-learn软件包实现。\n通过调整输入算法的最佳输入变量，各算法最优超参数通过格网化搜索方法获得。基于RFR算法精度较高：不确定性为（ϵ）为21.92%，偏差（β）为4.29%，斜率为0.64，均方根对数误差为0.235。\n</p>",
    "ds_quality": "<p> 采样随机森林回归（RFR）方法构建水质参数模型。随机森林是一种基本单元为决策树的集成学习方法，在训练过程中构建大量相互独立的决策树形成“随机森林”，最后综合这些决策树结果提高模型精度（例如输出平均值）。通过调整输入确定各水质参数算法的最佳输入变量，各算法最优超参数通过格网化搜索方法获得。利用大量野外调查和星地同步数据进行模型研究，结果表明RFR透明度算法精度较高，不确定性为（ϵ）为21.92%，偏差（β）为4.29%，斜率为0.64，均方根对数误差为0.235。</p>",
    "ds_acq_start_time": null,
    "ds_acq_end_time": null,
    "ds_acq_place": "",
    "ds_acq_lon_east": 96.76639,
    "ds_acq_lat_south": 33.648613,
    "ds_acq_lon_west": 72.78667,
    "ds_acq_lat_north": 49.8925,
    "ds_acq_alt_low": null,
    "ds_acq_alt_high": null,
    "ds_share_type": "login-access",
    "ds_total_size": 1547811693,
    "ds_files_count": 20,
    "ds_format": "TIF格式",
    "ds_space_res": "10米",
    "ds_time_res": "月",
    "ds_coordinate": "CGCS2000",
    "ds_projection": "Albers投影",
    "ds_thumbnail": "d840d65e-d096-41cc-ba76-32e304a1b2b8.png",
    "ds_thumb_from": 0,
    "ds_ref_way": "",
    "paper_ref_way": "",
    "ds_ref_instruction": "数据来源引用：新疆10m水体透明度数据集（2024年4-10月）来源于第三次新疆综合科学考察专项 \"空天地网一体化综合科考监测体系建设(2021xjkk1400)\"",
    "ds_from_station": null,
    "organization_id": "a5877b42-96ea-4f13-af7e-246f355413d6",
    "doi_value": "",
    "subject_codes": [
        "170.45"
    ],
    "quality_level": 1,
    "publish_time": "2025-12-04 18:16:20",
    "first_publish_time": null,
    "last_updated": "2026-01-14 10:57:52",
    "protected": false,
    "protected_to": null,
    "lang": "zh",
    "cstr": "33110.11.ariddc.01345",
    "license": null,
    "extra": null,
    "files_shape": [
        {
            "name": "2021xjkk1400-138-2024120801",
            "size": null,
            "is_dir": true
        }
    ],
    "features": null,
    "data_level": 0,
    "i18n": {
        "en": {
            "title": "Xinjiang 10-meter water body transparency dataset (April to October 2024)",
            "ds_abstract": "<p>This data uses 2024 Sentinel-2 optical satellite data as the data source, and a sampling random forest regression (RFR) method is used to build a remote sensing estimation model for lake water transparency in Xinjiang. The data is Abers projection in the CGCS2000 coordinate system with an accuracy of 10 meters. The numerical coefficient is 0.001, that is, the data pixel value is multiplied by 0.001 to obtain the actual water body transparency (m). </p>",
            "ds_source": "<p>Using 2024 Sentinel 2 optical satellite data L1C level data, data can be obtained from ESA's Copernicus Data Open Center (https://dataspace.copernicus.eu/browser/). </p>",
            "ds_process_way": "<p> Use Sentinel 2 A/B MSI with good radiation properties. Sentinel 2A/B MSI has 13 spectral bands with spatial resolutions of 10 m, 20 m and 60 m respectively. The revisit period is 10 days and 5 days after the double satellite network. It has demonstrated good performance in inland water bodies. Obtained L1C data from ESA's Copernicus Data Open Center.\nThe POLYMER algorithm is used for atmospheric correction. Further removal of the effects of sky light, solar flares and residual aerosol scattering is then needed:\nR_rs (λ)=(R(λ)-min⁡(R(865),R(2202)))/π\nWhere, Rrs is the remote-sensing reflectance of the water body, and R is the surface reflectance obtained by POLYMER correction.\nA sampling random forest regression (RFR) method was used to build a water quality parameter model. Random forest is an integrated learning method whose basic unit is a decision tree. During the training process, a large number of independent decision trees are built to form a \"random forest\", and finally the results of these decision trees are synthesized to improve the accuracy of the model (such as output average). The \"randomness\" of the RFR algorithm is mainly reflected in two aspects: when building each decision tree, the Bagging method (i.e., the self-service sampling method, where samples are put back after each sampling) is randomly sampled from the original training data set to generate a training data subset; When nodes are split, all feature variables are not used to participate in the comparison, but a subset of the feature variables is randomly selected to participate in node splitting. Two \"randomness\" make the RFR algorithm less prone to overfitting and has good tolerance for outliers and noise.\nThe most important hyperparameters of the RFR algorithm during the training process are: (1) The number of decision trees (n_estimators). The larger the n_estimators, the better the model results, but the longer the calculation time is; when a certain amount of data is reached, the model tends to be stable. (2) The maximum number of features (max_features) when the node is split. The decision tree finds the best split feature from randomly selected max_features when the node is split. (3) The maximum depth of the decision tree (max_depth). If it is not set (i.e. None), the decision tree will grow to the maximum extent until the segmentation termination condition is met. RFR is implemented through the Python scikit-learn software package.\nBy adjusting the optimal input variables of the input algorithm, the optimal hyperparameters of each algorithm are obtained through the grid search method. The algorithm based on RFR-has high accuracy: the uncertainty is 21.92%, the deviation (β) is 4.29%, the slope is 0.64, and the root mean square logarithmic error is 0.235.\n</p>",
            "ds_quality": "<p> A sampling random forest regression (RFR) method was used to build a water quality parameter model. Random forest is an integrated learning method whose basic unit is a decision tree. During the training process, a large number of independent decision trees are built to form a \"random forest\", and finally the results of these decision trees are synthesized to improve the accuracy of the model (such as output average). The optimal input variables for each water quality parameter algorithm are determined by adjusting the inputs, and the optimal hyperparameters of each algorithm are obtained through a grid search method. A large number of field surveys and satellite-earth synchronous data are used to conduct model research. The results show that the RFR transparency algorithm has high accuracy, with an uncertainty of 21.92%, a deviation of 4.29%, a slope of 0.64, and a root-mean-square logarithmic error of 0.235. </p>",
            "ds_ref_instruction": "Data source citation: Xinjiang's 10-meter water transparency dataset (April to October 2024) comes from the third Xinjiang comprehensive scientific expedition special project \"Construction of Integrated Comprehensive Scientific Research Monitoring System of Air, Space, Space and Network (2021 xjkk1400)\"",
            "ds_format": "TIF format",
            "ds_projection": "Albers projection",
            "ds_space_res": "10 meters",
            "ds_time_res": "months"
        }
    },
    "license_type": null,
    "doi_reg_from": "reg_local",
    "cstr_reg_from": "reg_local",
    "doi_not_reg_reason": null,
    "cstr_not_reg_reason": null,
    "is_paper_in_submitting": false,
    "ds_topic_tags": [
        "水体透明度"
    ],
    "ds_subject_tags": [
        "地理学"
    ],
    "ds_class_tags": [],
    "ds_locus_tags": [
        "新疆"
    ],
    "ds_time_tags": [
        2024
    ],
    "ds_contributors": [
        "刘铁",
        "段洪涛"
    ],
    "ds_meta_authors": [
        "段洪涛"
    ],
    "ds_managers": [
        "李锦"
    ],
    "category": "哨兵"
}