概覽
CPOS 資訊科技部營運高性能計算集群設施 (HPCF),配備 48 個節點、6,848 個 CPU 核心、124 張 GPU 卡、54 TB 記憶體以及 PB 級全快閃記憶體儲存裝置,並於 400GB 高速網路上運行。
目標用戶
- 需要運算服務以進行泛組學 (PanorOmics) 研究的香港大學教職員與學生
服務時間
- 於一般情況下可提供 24×7 全天候服務,惟系統進行維護期間提供有限度服務
提供泛組學科研服務
龐大的軟件庫
在過去 12 個月內完成
第三代高性能計算集群設施 (HPCF3)
🖥️ HPCF3 GPU 免費試用推廣 – 現已開放使用!
點擊 按此 了解更多詳情。
🖥️ HPCF3 Web Portal 網路入口網站 – 現已推出 BETA 測試版!
點擊 按此 了解更多詳情。
硬件與存取
🖥️HPCF3 集群概覽
HPCF3 登入節點可透過 SSH 存取,主機名稱為:
hpcf3.cpos.hku.hk
點擊 按此 了解更多詳情。
🧰作業系統與軟件
所有 HPCF3 伺服器均運行於 Rocky Linux 8 作業系統。
廣泛使用的生物資訊學軟件已預先安裝,所有用戶均可使用。
⚙️ 服務
用戶透過集群的作業排程系統 OpenPBS 提交工作。作業僅透過 PBS 工作在運算節點上執行。
➡️ 注意:系統嚴禁直接登入運算節點。
👤用戶帳號設定
如欲申請帳戶,請聯絡 itsupport.cpos@hku.hk.
- 所有帳戶申請均須經 CPOS 審批及批准。
- 獲批准的用戶將獲得一個 HPCF3 集群的 Linux 帳戶。
- 每個帳戶預設提供 100 GB 儲存空間;如需額外配額,可另行提出申請,惟須視乎資源供應情況及相關收費而定。
- 每個帳戶僅供獲授權的 教職員或學生本人 使用—— 嚴禁與他人共用帳戶。
- 雖然儲存系統已具備硬件冗餘保護機制,CPOS 仍 強烈建議 用戶定期自行 備份資料。.
- 用戶應熟悉 Linux/UNIX 作業環境,並具備基本系統操作知識。ITS 已提供 參考指南 供用戶查閱及參考。
🔐集群登入
用戶可透過 SSH 用戶端( PuTTY)進行遠端連線,網址如下: http://www.chiark.greenend.org.uk/~sgtatham/putty/
連線詳情
若要連線到 HPCF3 的登入節點,請使用 SSH 用戶端,開啟 SSH 終端會話,命令如下:
Hostname: hpcf3.cpos.hku.hk Username:
🔒網路要求
基於安全考量,HPCF3 伺服器僅限從 香港大學網路 內部存取。
-
校外用戶必須先透過 HKUVPN連線。詳情請參閱:
https://its.hku.hk/services/network-connectivity/hkuvpn/
📦資料傳輸
🪟 Windows 用戶
Windows 用戶可使用多款免費且成熟的圖形介面檔案傳輸工具,例如 WinSCP 或 FileZilla。透過這些工具,用戶只需以拖放方式,即可在集群前端節點與本機電腦之間輕鬆傳輸檔案。
🍏 Mac / 🐧 Linux 用戶
對於 Mac 或 Linux 用戶,命令列檔案傳輸工具 scp 通常已隨作業系統預先安裝。用戶可透過本機電腦的終端機(Terminal)連接至集群並傳輸檔案。
例如,假設用戶的 Mac 電腦上有一個名為 myfile 的本機檔案,並希望將其複製至集群帳戶的主目錄(Home Directory),可使用以下指令:
例子:
scp myfile userid@hpcf3.cpos.hku.hk:mydir/
此指令會將 myfile 複製至您 HPCF3 帳戶主目錄下的 mydir/ 目錄。
與其他以檔案或目錄作為參數的 Unix 指令一樣,來源及目的地路徑均可使用絕對路徑(以斜線 / 開頭)或相對路徑(不以斜線 / 開頭)來指定。
對於本機檔案或目錄,相對路徑是以目前工作目錄(Current Working Directory)為基準;而對於遠端檔案或目錄,相對路徑則是以遠端主機上的使用者主目錄(Home Directory)為基準。
📡大量資料傳輸指南
為維持校園網路服務的穩定性,大量資料傳輸請遵循以下指引:
- 超過 30 GB 的對外資料傳輸應安排於非辦公時間進行。
- 如預計傳輸量超過 500 GB,請提前通知 itsupport.cpos@hku.hk,以便我們與 ITS 協調相關安排。
⚠️未事先通知的大型資料傳輸,可能導致 ITS 暫時封鎖集群伺服器的對外網際網路連線。
收費
點擊 按此 了解更多詳情。
收費政策
👤HPCF 帳戶
- 每個 HPCF 用戶帳戶均按月收費,以支付集群服務的日常支援及維護成本。
- 計費週期由 每月 1 日開始。.
- 新增帳戶將由下一個計費週期起開始收費。例如,若新帳戶於某月 28 日完成開通並以電郵通知用戶,則當月餘下日期不另行收費,而於下一個月 1 日起開始計費。
- 所有 HPCF 用戶帳戶均須至少使用及繳付一個月費用,不設短期使用安排。
- 如需刪除用戶帳戶,相關申請將於當前計費週期的最後一天生效;在帳戶正式刪除前,仍須按正常標準收費。
📦用戶主目錄資料儲存空間
- 用戶主目錄儲存空間(/home/)按月收費,費用根據獲分配的磁碟配額計算,收費單位為每 100 GB(1 KB = 1000 Bytes)。
- 計費週期由 每月 1 日開始。.
- 新用戶帳戶的主目錄儲存空間將與帳戶月費一併由下一個計費週期起開始收費。
- 除非另有指定,所有新用戶帳戶均預設獲分配 100 GB 磁碟配額 。
- 如需增加磁碟配額,用戶須提交申請並抄送其 PI;獲批准後,調整後的配額將即時生效,並由當前計費週期起按新配額計費。
- 如需減少磁碟配額,用戶可提出申請,相關調整將於當前計費週期的最後一天生效;在配額正式下調前,仍按原有配額標準收費。
📦用戶群組資料儲存空間
- 用戶群組資料儲存空間(/home/groups/)按月收費,費用根據獲分配的磁碟配額計算,收費單位為每 100 GB(1 KB = 1000 Bytes)。
- 計費週期由 每月 1 日開始。.
- 如需增加磁碟配額,群組協調人須提交申請並抄送其 PI;獲批准後,調整後的配額將即時生效,並由當前計費週期起按新配額計費。
- 如需減少磁碟配額,用戶可提出申請,相關調整將於當前計費週期的最後一天生效;在配額正式下調前,仍按原有配額標準收費。
共置服務(Co-location Service)
- 對於新增的共置設備單元,每月支援服務費將於 完成 UAT 驗收後的下一個計費週期開始收取。.
重要備註
CPOS 保留根據實際資源使用情況及營運成本調整收費標準的權利。
就共置設備而言,所有香港大學財務處資產將由 CPOS 統一持有及管理。共置伺服器可能會整合至 HPCF 集群,並配置為運算節點使用;CPOS 將為共置伺服器上的工作執行設立專用工作佇列(Job Queue)。
請注意,CPOS IT 團隊負責集中管理 HPCF 內所有系統資源,包括共置設備。為確保港大研究社群能有效運用資源及提升整體效益,CPOS 可因應實際運作需要,將閒置的共置資源調配作工作運算之用,而毋須另行通知。
環境模組
HPCF3 集群採用 環境模組 系統管理集中安裝的軟件。系統已預先安裝多個版本的常用生物資訊學工具,方便用戶根據需要選擇及切換不同的軟件環境。
顯示可用模組
環境模組讓用戶能透過簡單指令載入、卸載及切換不同的軟件環境。
如欲查閱 HPCF3 上目前可用的軟件套裝及版本,請執行:
module avail
此指令會列出系統中所有已安裝並可供使用的模組。用戶可根據需要載入相應的軟件版本,以建立所需的執行環境。

即將推出。敬請期待。
HPCF3 集群的工作排程系統 (OpenPBS)
所有工作均須透過 OpenPBS 工作排程系統提交。提交工作時,用戶需指定所需資源,包括排程、CPU 核心數、記憶體容量及執行時間。當所需資源可用時,OpenPBS 會自動將工作安排至合適的運算節點執行,並按照系統的資源配額及使用限制進行管理。
點擊 按此 了解更多詳情。
第二代高性能計算集群設施 (HPCF2)
⚠️ HPCF2 將於 2026 年 10 月停止服務,現已停止接受新申請。請使用 HPCF3 提交新的服務及資源申請。
硬件與存取
HPCF2 集群
HPCF2 的主控節點(Master Node)為 Omics,用戶可透過 SSH 並使用下列主機名稱進行存取。
hpcf2.cpos.hku.hk.
HPCF2 集群由 8 個運算節點(Compute Nodes) 組成,其硬體規格如下:
| 伺服器名稱 | CPU 型號 | CPU 數量 | 每顆 CPU 核心數目 | 每台伺服器核心總數 | 每台伺服器執行緒總數 | 記憶體 (GB) |
| hpch01 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch02 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch03 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch04 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch05 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch06 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch07 | Intel Xeon E5-2650 v4 2.2GHz | 2 | 12 | 24 | 48 | 256 |
| hpch08 | Intel Xeon E5-2683 v4 2.1GHz | 2 | 16 | 32 | 64 | 512 |
| Total: | 200 | 400 | 2,304 |
作業系統與系統軟件
HPCF2 採用 CentOS 7 Linux 作業系統,並已預先安裝多款常用生物資訊學軟件,供用戶直接使用。
集群節點(Cluster Nodes)
主控節點(Master Nodes)
服務項目
- 透過 SSH 登入,以互動模式編譯及執行命令列程式。
- 向工作排程系統提交批次作業。
- 使用 SFTP 工具(如 FileZilla 或 WinSCP)進行檔案上傳及下載。
控制措施(Control Measures)
主控節點僅供上述用途使用。所有分析及運算工作應透過工作排程系統提交,並於運算節點(Compute Nodes)上執行。為確保系統穩定運作,主控節點不得用於執行大量消耗 CPU、記憶體或長時間運行的工作。
主控節點設有以下資源使用限制:
- 每位用戶於任何時間可使用的 CPU 資源總量上限為 400%(相當於 4 個 CPU 核心全速運行)。
- 每個工作(Job)的 最長執行時間為 10 分鐘。
- 每位用戶於任何時間可使用的 記憶體總量上限為 10 GB。
系統會持續監察各用戶於主控節點上的資源使用情況。當上述限制被超出時,系統將按資源使用量由高至低終止相關程序,直至資源使用恢復至允許範圍內。受影響用戶將自動收到電郵通知,列明被終止的程序詳情。
運算節點(Compute Nodes)
服務項目
- 用戶可於主控節點(Master Node)透過工作排程系統提交 PBS 工作至運算節點執行。 基於系統管理及資源調度需要,用戶不可直接登入運算節點執行工作。
帳戶申請與使用
- 如需申請 HPCF2 用戶帳戶,請聯絡 itsupport.cpos@hku.hk 。
- 所有帳戶申請均須經 CPOS 審批。
- 帳戶獲批後,系統將為用戶建立專屬 Linux 帳戶,以存取集群資源。
- 每個新帳戶預設提供 100 GB 磁碟配額。如有需要,可申請額外儲存空間,惟須視乎資源供應情況及相關收費安排而定。
- 每個帳戶僅限指定用戶本人使用,並由該用戶承擔相關責任。 嚴禁共用帳戶。
- 額外磁碟配額可按需要申請,並須視乎資源供應情況及相關收費安排而定。
- 雖然儲存系統設有硬體冗餘保護機制,但 CPOS 仍強烈建議 用戶定期將重要資料備份 至本地儲存裝置,以策安全。
- 用戶應具備基本 Linux/UNIX 使用知識。ITS 提供的 UNIX 使用指南可供參考: 按此 查看。
登入集群
用戶可透過 SSH(Secure Shell) 遠端連接至 HPCF2 的命令列環境。Windows 用戶可使用 PuTTY 等 SSH 用戶端,而 macOS 及 Linux 用戶則可直接使用系統內建的 SSH 指令。(如: http://www.chiark.greenend.org.uk/~sgtatham/putty/).
如需連接至 HPCF2 主控節點(Master Node),請使用 SSH 用戶端並以下列資料建立連線:
hostname: hpcf2.cpos.hku.hk
登入時請使用獲分配的 HPCF2 用戶帳戶名稱及密碼進行身份驗證。
注意: 基於資訊安全考慮,HPCF 伺服器僅接受來自港大校園網絡(HKU Network)的連線。如從校外網絡存取,必須先連接 HKU VPN,然後方可登入 HPCF 伺服器。 有關 HKU VPN 的設定及使用方法,請參閱 ITS 網頁: https://its.hku.hk/services/network-connectivity/hkuvpn/
數據傳輸
Windows 用戶可使用多款免費且成熟的圖形化安全檔案傳輸工具,例如 WinSCP 或 FileZilla (http://filezilla-project.org/)。透過這些工具,您可在本機電腦與 HPCF2 主控節點之間直接以拖放方式上傳及下載檔案。
macOS 及 Linux 作業系統一般已預先安裝 scp(Secure Copy) 指令,可透過終端機(Terminal)進行安全檔案傳輸。
例如,若您希望將本機上的檔案 myfile 複製到 HPCF2 主目錄下的 mydir/ 目錄,可執行:
scp myfile userid@hpcf2.cpos.hku.hk:mydir/
上述指令會將本機檔案 myfile 複製至 HPCF2 帳戶主目錄下的 mydir/ 目錄,並保留原有檔案名稱。
與其他 Unix/Linux 指令一樣,scp 的來源及目的地路徑均可使用:
絕對路徑(Absolute Path):以 / 開頭的完整路徑。
相對路徑(Relative Path):不以 / 開頭的路徑。
對於本機檔案,相對路徑是以目前工作目錄(Current Working Directory)為基準;對於 HPCF2 上的檔案,相對路徑則是以用戶的主目錄(Home Directory)為基準。
大量資料傳輸
為避免影響港大校園網絡的整體效能,ITS 建議所有涉及港大網絡以外地點的大量資料傳輸,應遵循以下安排:
傳輸量超過 30 GB 且需往返港大網絡以外地點的資料,應安排於非辦公時間進行。
如預計傳輸量超過 500 GB,請提前通知 itsupport.cpos@hku.hk,以便 CPOS 與 ITS 協調相關網絡流量安排。
如未有就大規模資料傳輸作出預先通知,ITS 可能會暫時限制 HPCF2 伺服器的對外網絡連線,以保障校園網絡服務的穩定性。
停止服務
HPCF2 將於 2026 年 10 月全面停止服務(End of Service)。
收費
點擊 按此 了解更多詳情。
收費政策
HPCF 帳戶
- 每個 HPCF 用戶帳戶均按月收費,以支付集群服務的日常支援及維護成本。
- 計費週期由每月 1 日開始。
- 新增帳戶將由下一個計費週期起開始收費。例如,若新帳戶於某月 28 日完成開通並以電郵通知用戶,則當月餘下日期不另行收費,而於下一個月 1 日起開始計費。
- 所有 HPCF 用戶帳戶均須至少使用及繳付一個月費用,不設短期使用安排。
- 如需刪除用戶帳戶,相關申請將於當前計費週期的最後一天生效;在帳戶正式刪除前,仍須按正常標準收費。
用戶主目錄資料儲存空間
- 用戶主目錄儲存空間(/home/)按月收費,費用根據獲分配的磁碟配額計算,收費單位為每 100 GB(1 KB = 1000 Bytes)。
- 計費週期由每月 1 日開始。
- 新用戶帳戶的主目錄儲存空間將與帳戶月費一併由下一個計費週期起開始收費。
- 除非另有說明,每個新用戶帳戶均預設提供 100 GB 磁碟配額。
- 如需增加磁碟配額,用戶須提交申請並抄送其 PI;獲批准後,調整後的配額將即時生效,並由當前計費週期起按新配額計費。
- 如需減少磁碟配額,用戶可提出申請,相關調整將於當前計費週期的最後一天生效;在配額正式下調前,仍按原有配額標準收費。
用戶群組資料儲存空間
- 用戶群組資料儲存空間(/home/groups/)按月收費,費用根據獲分配的磁碟配額計算,收費單位為每 100 GB(1 KB = 1000 Bytes)。
- 計費週期由每月 1 日開始。
- 如需增加磁碟配額,群組協調人須提交申請並抄送其 PI;獲批准後,調整後的配額將即時生效,並由當前計費週期起按新配額計費。
- 如需減少磁碟配額,用戶可提出申請,相關調整將於當前計費週期的最後一天生效;在配額正式下調前,仍按原有配額標準收費。
共置服務(Co-location Service)
- 對於新增的共置設備單元,每月支援服務費將於完成 UAT 驗收後的下一個計費週期開始收取。
重要備註
CPOS may revise the charges from time to time based on actual usages and recovery needs.
就共置設備而言,所有香港大學財務處資產將由 CPOS 統一持有及管理。共置伺服器可能會整合至 HPCF 集群,並配置為運算節點使用;CPOS 將為共置伺服器上的工作執行設立專用工作佇列(Job Queue)。
請注意,CPOS IT 團隊負責集中管理 HPCF 內所有系統資源,包括共置設備。為確保港大研究社群能有效運用資源及提升整體效益,CPOS 可因應實際運作需要,將閒置的共置資源調配作工作運算之用,而毋須另行通知。
環境模組
HPCF2 採用 環境模組(Environment Modules) 系統管理集中安裝的軟體。系統已預先安裝多個版本的常用生物資訊學軟體,方便用戶按需要選擇及切換不同的軟體環境。請注意:環境模組功能僅適用於新版 HPCF2 集群,不適用於舊版 HPCF(statgenpro)。
可用模塊
環境模組讓用戶能透過簡單指令載入、卸載及切換不同的軟體環境。 如欲查閱集群上目前可用的軟體套件及版本,請執行以下指令:
[itsupport@omics ~]$module avail --------------------------------------------- /software/Modules/modulefiles ---------------------------------------------- ANNOVAR/2017Jul16 HTSeq/0.9.1 (D) STAR/2.5.2a bwa/0.7.12 BEDTools/2.12.0 MACS/2.0.10-2012.06.06 STAR/2.5.3a (D) bwa/0.7.17 (D) BEDTools/2.17.0 MACS/2.1.0-2015.04.20 (D) TrimGalore/0.4.1 cutadapt/1.8.1 BEDTools/2.27.1 (D) MUMmer/3.22 TrimGalore/0.4.5 (D) cutadapt/1.15 (D) BioPerl/1.7.2 MUMmer/3.23 (D) Trimmomatic/0.33 idba/1.1.3 Canu/1.5 NCBI-blast/2.2.27+ Trimmomatic/0.36 (D) java/7.0_25 Canu/1.6 (D) NCBI-blast/2.7.1+ (D) VerifyBamID/1.1.2 java/7.0_80 CellRanger/2.0.1 Oncotator/1.9.6.1 VerifyBamID/1.1.3 (D) java/8.0_161 (D) CellRanger/2.1.0 (D) PEAR/0.9.10 bamUtil/1.0.13 java/9.0.4 DESeq2/1.10.1 PEAR/0.9.11 (D) bamUtil/1.0.14 (D) miniconda2/4.3.31 DESeq2/1.18.1 (D) Perl/5.26.1 bamtools/2.3.0 muTect/1.1.4 EBSeq/1.9.3 Picard/2.0.1 bamtools/2.5.1 (D) muTect/1.1.5 (D) EBSeq/1.18 (D) Picard/2.17.4 (D) bcl2fastq/2.19 python2/2.7.14 FASTX-toolkit/0.0.13.2 QIIME/1.9.1 bcl2fastq/2.20 (D) python3/3.6.4 FASTX-toolkit/0.0.14 (D) QIIME2/2017.12 bedGraphToBigWig/4 samtools/0.1.18 FastQC/0.11.2 R/3.2.5 bismark/0.14.3 samtools/1.3 FastQC/0.11.7 (D) R/3.4.3 (D) bismark/0.19.0 (D) samtools/1.6 (D) GenomeAnalysisTK/3.5 RNAmmer/1.2 bowtie/1.0.0 strelka/1.0.15 GenomeAnalysisTK/3.7 RSEM/1.2.31 bowtie/1.2.2 (D) strelka/2.8.4 (D) GenomeAnalysisTK/3.8 (D) RSEM/1.3.0 (D) bowtie2/2.2.5 tRNAscan-SE/1.3.1 HOMER/4.9 SPAdes/3.10.0 bowtie2/2.3.4 (D) HTSeq/0.6.1 SPAdes/3.11.1 (D) bwa/0.6.2 -------------------------------------- /opt/Lmod/7.7.14/lmod/lmod/modulefiles/Core --------------------------------------- lmod/7.7.14 settarg/7.7.14 Where: D: Default Module
Load module
Then you can execute these centrally installed software package by loading the corresponding module(s). For example, if you would like to run the BWA alignment tool, you can use the command “module load” to enable the bwa module environment:
[itsupport@omics ~]$ module load bwa bwa/0.7.17 is loaded
Then, the default version of bwa module would be loaded. You can then use the bwa tool for your subsequent data analysis:
[itsupport@omics ~]$ bwa Program: bwa (alignment via Burrows-Wheeler transformation) Version: 0.7.17-r1188 Contact: Heng LiUsage: bwa [options] Command: index index sequences in the FASTA format mem BWA-MEM algorithm fastmap identify super-maximal exact matches ...
List loaded modules
To list out those modules that are currently loaded, use the command “module list”:
[itsupport@omics ~]$ module list Currently Loaded Modules: 1) bwa/0.7.17
Unload modules
When you no longer need to execute that software, use the command “module unload” to unload it from your current session:
[itsupport@omics ~]$ module unload bwa bwa/0.7.17 is unloaded
Search for available versions of a module
If you would like to use another version but not the default one, you can search and specify a particular version of the software module:
itsupport@omics ~]$module avail bwa
------------------------------- /software/Modules/modulefiles -------------------------------
bwa/0.6.2 bwa/0.7.12 bwa/0.7.17 (D)
Where:
D: Default Module
Load specific version of a module
Load an older version (0.7.12) of bwa:
[itsupport@omics ~]$ module load bwa/0.7.12 bwa/0.7.12 is loaded
You can then execute this older version of bwa now:
[itsupport@omics ~]$ bwa Program: bwa (alignment via Burrows-Wheeler transformation) Version: 0.7.12-r1039 Contact: Heng LiUsage: bwa [options] Command: index index sequences in the FASTA format mem BWA-MEM algorithm ...
Loading alternative version of a module
[itsupport@omics ~]$ module load bowtie2/2.2.5 bowtie2/2.2.5 is loaded [itsupport@omics ~]$ module load bowtie2/2.3.4 bowtie2/2.2.5 is unloaded bowtie2/2.3.4 is loaded The following have been reloaded with a version change: 1) bowtie2/2.2.5 => bowtie2/2.3.4
Quick reference for Module Commands
| Command | 功能說明 |
| module list | List currently loaded module(s) |
| module avail | Show what modules are available for loading |
| module avail [name] | Show only the modules that are available for the application named [name] |
| module keyword [word1] [word2] … | Show available modules matching the search criteria |
| module whatis [module_name] | Show description of particular module |
| module help [module_name] | Show help information |
| module load [module_name] | Configure your environment according to modulefile(s) |
| module load [module_name]/[version] | Load specific version of a module |
| module load [mod A] [mod B] … | Load a list of modules |
| module unload [module_name] | Roll back configuration performed by the modulefile(s) |
| module unload [mod A] [mod B] … | Unload a list of modules |
| module swap [module A] [module B] | Unload modulefile A and load modulefile B |
| module purge | Unload all modules currently loaded |
Using module in shell scripts
Here is an example of using module in the shell script:
#!/bin/bash # cleanup first module purge # load the modules we need for this script module load bwa/0.7.12 # perform the data analysis bwa mem reference.fa reads1.fq reads2.fq > aligned_pairs.sam
Loading modules with prerequisites
Some modules may depend on other modules in order to work properly and user will be prompted to load the prerequisite modules when loading such modules.
[itsupport@omics ~]$module load bismark Lmod has detected the following error: Cannot load module "bismark/0.19.0" without these module(s) loaded: bowtie2 While processing the following module(s): Module fullname Module Filename --------------- --------------- bismark/0.19.0 /software/Modules/modulefiles/bismark/0.19.0.lua
[itsupport@omics ~]$module load bowtie2 bowtie2/2.3.4 is loaded [itsupport@omics ~]$module load bismark bismark/0.19.0 is loaded
軟件
HPCF2 已集中安裝多款常用科研及生物資訊學軟件。用戶可透過環境模組系統載入所需軟件,無需自行安裝。以下為目前可供使用的軟件清單:
| 軟件名稱 | HPCF2 模組名稱 | 已安裝版本 | 官方網站 | 功能簡介 |
| 7-Zip | 7-Zip | 16.02 | http://www.7-zip.org/ | file archiver used to place groups of files within compressed containers known as “archives”. |
| ANNOVAR | ANNOVAR | 2020Jun08 | http://annovar.openbioinformatics.org/ | annotate genetic variants detected from diverse genomes |
| BamTools | bamtools | 2.3.0 2.5.1 | https://github.com/pezmaster31/bamtools | end-user’s toolkit for handling BAM files |
| BamUtil | bamUtil | 1.0.13 1.0.14 | https://genome.sph.umich.edu/wiki/BamUtil | end-user’s programs for operating BAM/SAM files |
| bcl2fastq | bcl2fastq | 2.19 2.20 | https://support.illumina.com/sequencing/sequencing_software/bcl2fastq-conversion-software.html | demultiplexes data and converts BCL files generated by Illumina sequencing systems to standard FASTQ file formats for downstream analysis |
| bedGraphToBigWig | bedGraphToBigWig | 4 | https://www.encodeproject.org/software/bedgraphtobigwig/ | Convert bedGraph to bigWig file |
| BEDTools | BEDTools | 2.12.0 2.17.0 2.27.1 | http://bedtools.readthedocs.io/en/latest/ | a collection of utility tools for a wide-range of genomics analysis tasks such as merging and shuffling genomic intervals from multiple files in widely-used genomic file formats such as BAM, BED, GFF/GTF, VCF |
| BioPerl | BioPerl | 1.7.2 | http://bioperl.org/ | open source Perl tools for bioinformatics, genomics and life science |
| Bismark | bismark | 0.14.3 0.19.0 0.20.0 | https://www.bioinformatics.babraham.ac.uk/projects/bismark/ | A tool to map bisulfite converted sequence reads and determine cytosine methylation states |
| Bowtie | bowtie | 1.0.0 1.2.2 2.2.5 2.3.4 2.3.4.1 2.3.4.3 2.4.2 | http://bowtie-bio.sourceforge.net/index.shtml | An ultrafast memory-efficient short read aligner |
| Bowtie 2 | bowtie2 | 2.2.5 2.3.4 2.3.4.1 2.3.4.3 2.4.2 | http://bowtie-bio.sourceforge.net/bowtie2/index.shtml | An ultrafast and memory-efficient tool for aligning sequencing reads to long reference sequences |
| BWA | bwa | 0.6.2 0.7.12 0.7.17 | http://bio-bwa.sourceforge.net/ | Alignment of short reads via Burrows-Wheeler transformation on an indexed reference sequence |
| Canu | Canu | 1.5 1.6 1.9 | http://canu.readthedocs.io/ | a single molecule sequence assembler for genomes large and small |
| CellRanger | CellRanger | 2.0.1 2.1.0 2.2.0 3.0.2 3.1.0 4.0.0 6.1.2 | https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/latest/what-is-cell-ranger | a set of analysis pipelines that process Chromium single-cell RNA-seq output to align reads, generate gene-cell matrices and perform clustering and gene expression analysis |
| cutadapt | cutadapt | 1.16 1.8.1 2.3 3.4 | https://github.com/marcelm/cutadapt | removes adapter sequences from high-throughput sequencing data |
| DESeq2 | DESeq2 | 1.10.1 1.18.1 | https://bioconductor.org/packages/release/bioc/html/DESeq2.html | Differential gene expression analysis based on the negative binomial distribution |
| EBSeq | EBSeq | 1.9.3 1.18 | http://bioconductor.org/packages/release/bioc/html/EBSeq.html | An R package for gene and isoform differential expression analysis of RNA-seq data |
| FastQC | FastQC | 0.11.8 0.11.9 | https://www.bioinformatics.babraham.ac.uk/projects/fastqc/ | A quality control tool for high throughput sequence data |
| FASTX-Toolkit | FASTX-toolkit | 0.0.14 1.3.0 | http://hannonlab.cshl.edu/fastx_toolkit/ | a collection of command line tools for Short-Reads FASTA/FASTQ files preprocessin |
| Genome Analysis Toolkit (GATK) | GenomeAnalysisTK | 3.7 3.8.1.0 4.1.9.0 4.2.0.0 | https://software.broadinstitute.org/gatk/ | a wide variety of tools with a primary focus on variant discovery and genotyping |
| HOMER | HOMER | 4.9 4.11 | http://homer.ucsd.edu/homer/ | HOMER (Hypergeometric Optimization of Motif EnRichment) is a suite of tools for Motif Discovery and next-gen sequencing analysis |
| HTSeq | HTSeq | 0.6.1 0.9.1 | https://htseq.readthedocs.io/en/release_0.9.1/ | a Python package that provides infrastructure to process data from high-throughput sequencing assay |
| IDBA | idba | 1.1.3 | https://github.com/loneknightpy/idba | basic iterative de Bruijn graph assembler for second-generation sequencing reads |
| Java | java | 8.0_161 9.0.4 10.0.2 11.0.9 12.0.2 13.0.2 | https://www.java.com | a general-purpose computer-programming language that is concurrent, class-based, object-oriented, and specifically designed to have as few implementation dependencies as possible |
| MACS | MACS | 2.0.10-2012.06.06 2.1.0-2015.04.20 2.1.1-2016.03.09 2.1.2-2019.09.06 | http://liulab.dfci.harvard.edu/MACS/ | Model-based Analysis of ChIP-Seq (MACS), for identifying transcript factor binding sites |
| MUMmer | MUMmer | 3.22 3.23 4.0.ob2 | http://mummer.sourceforge.net/ | a system for rapidly aligning entire genomes, whether in complete or draft form |
| muTect | muTect | 1.1.4 1.1.5 | http://archive.broadinstitute.org/cancer/cga/mutect | identification of somatic point mutations in next generation sequencing data of cancer genomes |
| NCBI-blast | NCBI-blast | 2.2.27+ 2.7.1+ | https://blast.ncbi.nlm.nih.gov/Blast.cgi | BLAST finds regions of similarity between biological sequences. The program compares nucleotide or protein sequences to sequence databases and calculates the statistical significance. |
| Oncotator | Oncotator | 1.9.6.1 1.9.9.0 | http://portals.broadinstitute.org/oncotator/ | a web application for annotating human genomic point mutations and indels with data relevant to cancer researchers |
| PEAR | PEAR | 0.9.10 0.9.11 | https://www.h-its.org/downloads/pear-academic/ | an ultrafast, memory-efficient and highly accurate pair-end read merger. |
| Perl | Perl | 5.26.1 | https://www.perl.org | a high-level, general-purpose, interpreted, dynamic programming languages |
| Picard | Picard | 2.17.4 2.18.9 2.25.2 | https://broadinstitute.github.io/picard/ | command-line utilities that manipulate SAM files |
| Python 2 | python2 | 2.7.14 | https://www.python.org | an interpreted high-level programming language for general-purpose programming |
| Python 3 | python3 | 3.6.4 3.7.10 3.9.2 | https://www.python.org | an interpreted high-level programming language for general-purpose programming |
| QIIME | QIIME | 1.9.1 | http://qiime.org/ | an open-source bioinformatics pipeline for performing microbiome analysis from raw DNA sequencing data |
| QIIME2 | QIIME2 | 2017.12 2019.7 | https://qiime2.org/ | a microbiome analysis package with a focus on data and analysis transparency |
| R | R | 3.4.3 3.5.1 3.6.1 4.1.0 | https://www.r-project.org | a programming language and free software environment for statistical computing and graphics |
| RNAmmer | RNAmmer | 1.2 | http://www.cbs.dtu.dk/cgi-bin/nph-sw_request?rnammer | predicting ribosomal RNA genes in full genome sequences |
| RSEM | RSEM | 1.2.31 1.3.0 1.3.3 | http://deweylab.github.io/RSEM/ | accurate quantification of gene and isoform expression from RNA-Seq data |
| SAMtools | samtools | 1.6 1.8 1.9 1.11 | samtools.sourceforge.net | SAM Tools provide various utilities for manipulating alignments in the SAM format, including sorting, merging, indexing and generating alignments in a per-position forma |
| SPAdes | SPAdes | 3.10.0 3.11.1 3.13.0 | http://cab.spbu.ru/software/spades/ | St. Petersburg genome assembler – is an assembly toolkit containing various assembly pipelines |
| STAR | STAR | 2.7.8a 2.7.9a | https://github.com/alexdobin/STAR | RNA-seq aligner |
| strelka | strelka | 1.0.15 2.8.4 2.9.7 | https://github.com/Illumina/strelka | germline and somatic small variant caller |
| TrimGalore | TrimGalore | 0.4.1 0.4.5 | https://www.bioinformatics.babraham.ac.uk/projects/trim_galore/ | A wrapper tool around Cutadapt and FastQC to consistently apply quality and adapter trimming to FastQ files |
| Trimmomatic | Trimmomatic | 0.33 0.36 0.38 | http://www.usadellab.org/cms/index.php?page=trimmomatic | A flexible read trimming tool for Illumina NGS data |
| tRNAscan-SE | tRNAscan-SE | 1.3.1 | http://eddylab.org/software.html | tRNA detection in large-scale genome sequence |
| VerifyBamID | VerifyBamID | 1.1.2 1.1.3 | https://genome.sph.umich.edu/wiki/VerifyBamID | verifies whether the reads in particular file match previously known genotypes for an individual (or group of individuals), and checks whether the reads are contaminated as a mixture of two samples |
註:上述清單並非完整列表。如需查詢特定軟件是否可用,請聯絡 CPOS。
HPCF2 工作排程系統(PBS Pro)
HPCF2 集群採用 PBS Pro 作為工作排程系統。所有運算工作均須透過批次作業方式提交,並指定所需資源,例如工作排程(Queue)、CPU 數量、記憶體容量及執行時間。系統會在資源可用時,自動安排工作於運算節點上執行,並受資源使用上限及工作排程規則所限制。
一般工作排程
| 排程名稱 | 最大 CPU 核心數 | 記憶體 (GB)使用上限 | 同時執行工作上限 | 排程工作上限 | 執行時數上限 |
| small | 2 | 10 | 18 | 40 | 6 |
| small_ext | 2 | 10 | 6 | 12 | 60 |
| medium | 12 | 50 | 12 | 25 | 24 |
| medium_ext | 12 | 50 | 6 | 8 | 60 |
| large | 12 | 120 | 3 | 4 | 84 |
| legacy | 12 | 45 | 8 | 16 | 96 |
| test | 24 | 190 | 1 | 1 | 1 |
特殊工作排程(需另行申請)
如您的工作需要的運算資源超出一般工作排程所提供的限制,請向 CPOS 提交工作執行計劃及所需資源詳情。
CPOS 將根據當前集群資源使用情況及整體需求進行評估。在可行的情況下,可為特定項目設立臨時專用工作排程(Custom Job Queue),以提供額外或專屬的運算資源,支援相關工作執行。
工作腳本(Job Scripting)
PBS 工作指令(PBS Job Directives)
以下列出部分常用的 PBS 工作指令(PBS Job Directives)。這些設定可寫入工作腳本(Job Script)中,亦可在使用 qsub 提交工作時於命令列指定。
| PBS 工作指令(PBS Job Directives) | 功能說明 |
| #PBS -A acct | Causes the job time to be charged to “acct”. |
| #PBS -N myJob | Assigns a job name. The default is the name of PBS job script. |
| #PBS -l nodes=4:ppn=2 | The number of nodes and processors per node. |
| #PBS -l walltime=01:00:00 | Sets the maximum wall-clock time during which this job can run. (walltime=hh:mm:ss) |
| #PBS -l mem=n | Sets the maximum amount of memory allocated to the job. |
| #PBS -q queuename | Assigns your job to a specific queue. |
| #PBS -o mypath/my.out | The path and file name for standard output. |
| #PBS -e mypath/my.err | The path and file name for standard error. |
| #PBS -j oe | Join option that merges the standard error stream with the standard output stream of the job. |
| #PBS -M email-address | Sends email notifications to a specific user email address. |
| #PBS -m | Set email to be sent to the user when: |
| #PBS -m a | a – the job aborts |
| #PBS -m b | b – the job begins |
| #PBS -m e | e – the job ends |
| #PBS -r n | Indicates that a job should not rerun if it fails. |
| #PBS -S shell | Sets the shell to use. Make sure the full path to the shell is correct. |
| #PBS -V | Exports all environment variables to the job. |
| #PBS -W | Used to set job dependencies between two or more jobs. |
以下列出部分常用的 PBS 工作指令,可用於設定工作執行環境及所需資源。
| PBS 指令 | 功能說明 |
| -S /bin/bash | Specifies which shell to use. |
| -N JobName | Gives a name to the job. The name will appear in the output of qstat. |
| -M userame@hku.hk | Specifies an email address where notification messages will be sent. |
| -m abe (or any subset of a, b, and e) | Specifying when an email will be sent: a — abort, b — begin, e — end. See notes on email below. |
PBS 工作輸出/錯誤檔案
您可自行指定 PBS 工作的標準輸出(Standard Output)及標準錯誤輸出(Standard Error)檔案名稱及儲存位置。
| PBS 工作指令(PBS Job Directives) | 功能說明 |
| #PBS -o mypath/my.out | The path and file name for standard output |
| #PBS -e mypath/my.err | The path and file name for standard error |
您亦可只指定輸出檔案的存放目錄,由 PBS 自動產生檔案名稱。
| PBS 工作指令(PBS Job Directives) | 功能說明 |
| #PBS -o mypath | The path for standard output. Output file will be generated as, e.g. 123.omics.OU |
| #PBS -e mypath | The path for standard error. Error file will be generated as, e.g. 123.omics.ER |
如未在工作腳本中指定 -o 或 -e 選項,PBS 將使用工作提交時的目前工作目錄(Working Directory)作為輸出位置,並自動產生預設的輸出及錯誤檔案名稱。
| Output / Error Files | 功能說明 |
| Output File | File with name $PBS_JOBNAME.o$PBS_JOBID would be generated, e.g. myfirstjob.o123 |
| Error File | File with name $PBS_JOBNAME.e$PBS_JOBID would be generated, e.g. myfirstjob.e123 |
PBS Pro 會以提交工作時所在的目錄作為工作目錄(Working Directory),而非工作腳本(Job Script)所在的位置。
請注意,PBS 排程系統預設會將相關輸出檔案儲存於您的主目錄(Home Directory)下。為避免混淆及方便管理, 建議在工作腳本中使用完整路徑(Full Path)指定輸出檔案位置及名稱。.
注意事項
在 Omics 集群所使用的新 PBS 排程系統中,PBS 工作變數(例如常用的 $PBS_JOBID 及 $PBS_JOBNAME)不會於 #PBS 指令內被自動展開(Resolve)。此行為與舊版 statgenpro 集群有所不同,編寫工作腳本時請特別留意。 NOT be resolved in the #PBS job directives at the Omics cluster with new PBS scheduling system (that is different from the statgenpro cluster).
互動式 PBS 工作(Interactive PBS Jobs)
PBS 不僅支援批次作業(Batch Jobs),亦支援互動式工作(Interactive Jobs)。透過互動式工作,用戶可直接於運算節點上執行指令及使用軟體,而無需先建立工作腳本。
例如,用戶可在運算節點上啟動 R、Python 或其他開發環境進行測試及除錯,並在獲分配的執行時間(Walltime)內進行互動操作。
提交互動式工作時,無需準備 PBS 工作腳本,只需於 qsub 指令中直接指定所需資源。
例如,以下 PBS 工作腳本:
#PBS -l nodes=1:ppn=4 #PBS -l mem=2gb #PBS -l walltime=15:00:00 #PBS -q small
可改為直接使用以下 qsub 指令提交:
qsub -I -q small -l nodes=1:ppn=4,walltime=15:00:00,mem=2gb
當符合指定資源需求的運算節點可供使用時,PBS 排程系統會自動為用戶分配所需資源(例如 4 個 CPU 核心),並將用戶登入至其中一個運算節點,以互動模式執行工作。
請注意,為避免資源閒置而未被有效使用,所有互動式 PBS 工作(即使用 qsub -I 提交的工作)如連續 30 分鐘沒有任何操作或活動,系統將自動終止工作並登出用戶,以釋放已分配的運算資源供其他用戶使用。
工作管理(Job Management)
提交批次工作(Batch Jobs)
批次工作(Batch Jobs)讓用戶將工作提交至工作排程系統,並由系統在所需資源可用時自動安排於集群的運算節點上執行。
用戶需透過工作腳本(Job Script)提交工作。工作腳本包含執行程式所需的指令及資源需求設定。工作完成後,執行結果及訊息會輸出至檔案,供日後查閱及分析。
一般而言,任何可於 Linux 命令列執行的指令或程式,都可寫入 PBS 工作腳本並透過批次方式執行。
以下為一個簡單的 PBS 工作腳本範例(檔案名稱:simple.sh):
#!/bin/bash #PBS -l nodes=1:ppn=1 #PBS -l mem=2g #PBS -l walltime=00:01:00 #PBS -m ae #PBS -N omics-simple #PBS -q small
module load bamtools/2.5.1 bamtools -v
完成工作腳本後,可使用以下指令提交工作:
$ qsub simple.sh
工作提交後,PBS 會根據所指定的資源需求及佇列狀況,於資源可用時自動安排工作在運算節點上執行。
如需設定工作名稱、所需資源、執行時間或輸出檔案等參數,請參閱前述 「PBS 工作指令(PBS Job Directives)」 章節。
檢查工作狀態(Check Job Status)
提交工作後,可使用 qstat 指令查看工作狀態及各工作佇列的使用情況。
當運算節點有足夠資源可供使用時,工作會進入 R(Running) 狀態,表示正在執行;如所需資源暫時不足,工作則會顯示為 Q(Queued) 狀態,表示正在佇列中等待執行。當工作完成後,狀態會顯示為 C(Completed)。
顯示執行中的工作
$ qstat -rn1
例子:
[itsupport@omics ~]$qstat -rn1
omics:
Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time
--------------- -------- -------- ---------- ------ --- --- ------ ----- - -----
1649.omics itsupport large BWA_19NT 121502 1 12 90gb 48:00 R hpch01/0*12
顯示排隊中/暫停中的工作
$ qstat -i
例子:
[itsupport@omics ~]$qstat -i
omics:
Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time
--------------- -------- -------- ---------- ------ --- --- ------ ----- - -----
1651.omics itsupport large BWA_19T --- 1 12 90gb 48:00 Q
顯示所有工作(包括已完成的工作)
$ qstat -xan1
例子:
[itsupport@omics Exome]$qstat -xan1
omics:
Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time
--------------- -------- -------- ---------- ------ --- --- ------ ----- - -----
1648.omics itsupport large BWA_23T 127476 1 12 45gb 48:00 F 00:03 hpch08/0*12
1649.omics itsupport large BWA_23NT 121502 1 12 90gb 48:00 R 01:12 hpch08/0*12
1650.omics itsupport large BWA_43T 5940 1 12 45gb 48:00 F 00:01 hpch08/0*12
1651.omics itsupport large BWA_43NT -- 1 12 90gb 48:00 Q -- --
Show details of a job
$ qstat -xf JobID
例子:
[itsupport@omics]$qstat -f 1649 Job Id: 1649.omics Job_Name = BWA_34NT Job_Owner = itsupportn@omics resources_used.cpupercent = 916 resources_used.cput = 05:56:56 resources_used.mem = 59554972kb resources_used.ncpus = 12 resources_used.vmem = 92494740kb resources_used.walltime = 01:28:38 job_state = R ...
檢查已完成工作的資源使用情況
功能: 顯示指定工作的詳細資訊及資源使用統計資料。
用法: myjob
例子:
$ myjob 440356 Job information and usage summary of your HPCF job 440356 : +-------------------+----------+-------+-------------+------+---------------------+-----------+------+-------+ | jobid | username | queue | jobname | E S | End Time | walltime% | mem% | cpu% | +-------------------+----------+-------+-------------+------+---------------------+-----------+------+-------+ | 440356.statgenpro | kelvin | large | Large | 0 | 2016-08-30 04:31:30 | 25.36 | 8.28 | 14.51 | +-------------------+----------+-------+-------------+------+---------------------+-----------+------+-------+ +---------------------+---------------------+----------+----------+------+------+-------+----------+--------+ | Submit Time | Start Time | wtime@ | wtime# | mem@ | mem# | vmem@ | CPUTime@ | nproc# | +---------------------+---------------------+----------+----------+------+------+-------+----------+--------+ | 2016-08-29 22:26:16 | 2016-08-29 22:26:18 | 06:05:13 | 24:00:00 | 3.31 | 40gb | 4.75 | 10:35:53 | 12 | +---------------------+---------------------+----------+----------+------+------+-------+----------+--------+ E S = Exit Status ; % = usage percentage; # = requested ; @ = used ; mem@/vmem@ in GB ; nproc = number of processors
只有工作擁有者(Job Owner)才可執行此查詢及查閱相關資源使用資訊。
$ myjob 123456 ERROR: You (kelvin) not the owner of the job 123456.
檢查已完成工作的資源使用摘要
功能: 顯示近期已完成工作的摘要資訊及資源使用統計。
用法: myjobs [-v] [-j] []
說明: 顯示近期已完成工作的摘要資訊及資源使用統計。
可選參數
-v verbose mode with resource usage data of walltime/memory/cpu requested and used
-j Last jobs by JobID numbers (instead of the default by End Time). Note that the jobs in the list by JobID may be different from that by End Time.
specifying the number of jobs (default:20) to display
例子:
$ myjobs 5 Your last 5 HPCF completed jobs (by End Time): +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+ | jobid | username | queue | jobname | E S | End Time | cpu% | mem% | wtime% | +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+ | 458952.statgenpro | kelvin | large | GATK_CK02 | 0 | 2016-10-28 16:45:05 | 69.82 | 10.80 | 0.42 | | 458860.statgenpro | kelvin | large | GATK_NG07 | 0 | 2016-10-28 11:19:56 | 53.44 | 10.70 | 0.40 | | 458853.statgenpro | kelvin | large | GATK_CK03 | 0 | 2016-10-28 11:13:53 | 69.20 | 10.80 | 0.42 | | 458862.statgenpro | kelvin | large | GATK_YJ01 | 0 | 2016-10-28 11:09:05 | 49.90 | 10.30 | 0.24 | | 458858.statgenpro | kelvin | large | GATK_KC04 | 0 | 2016-10-28 10:57:39 | 55.96 | 10.30 | 0.11 | +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+ E S = Exit Status ; % = usage percentage; wtime = walltime
詳細模式
$ myjobs -v 5 Your last 5 HPCF completed jobs (by End Time): +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+----------+--------+-------+-------+----------+-----------+ | jobid | username | queue | jobname | E S | End Time | cpu% | mem% | wtime% | CPUTime@ | CPUno# | mem@ | mem# | wtime@ | wtime# | +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+----------+--------+-------+-------+----------+-----------+ | 458952.statgenpro | kelvin | large | GATK_CK02 | 0 | 2016-10-28 16:45:05 | 69.82 | 10.80 | 0.42 | 00:41:55 | 2 | 1.08 | 10.00 | 00:30:01 | 120:00:00 | | 458860.statgenpro | kelvin | large | GATK_NG07 | 0 | 2016-10-28 11:19:56 | 53.44 | 10.70 | 0.40 | 00:31:03 | 2 | 1.07 | 10.00 | 00:29:03 | 120:00:00 | | 458853.statgenpro | kelvin | large | GATK_KC03 | 0 | 2016-10-28 11:13:53 | 69.20 | 10.80 | 0.42 | 00:42:17 | 2 | 1.08 | 10.00 | 00:30:33 | 120:00:00 | | 458862.statgenpro | kelvin | large | GATK_YJ01 | 0 | 2016-10-28 11:09:05 | 49.90 | 10.30 | 0.24 | 00:16:57 | 2 | 1.03 | 10.00 | 00:16:59 | 120:00:00 | | 458858.statgenpro | kelvin | large | GATK_KC04 | 0 | 2016-10-28 10:57:39 | 55.96 | 10.30 | 0.11 | 00:08:55 | 2 | 1.03 | 10.00 | 00:07:58 | 120:00:00 | +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+----------+--------+-------+-------+----------+-----------+ E S = Exit Status ; % = usage percentage ; wtime = walltime ; @ = used ; # = requested ; CPUno = number of processors ; mem@/vmem@/mem# in GB
依 Job ID 排序顯示最近的工作
$ myjobs -j 5 Your last 5 HPCF completed jobs (by JobID): +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+ | jobid | username | queue | jobname | E S | End Time | cpu% | mem% | wtime% | +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+ | 458952.statgenpro | kelvin | large | GATK_CK02 | 0 | 2016-10-28 16:45:05 | 69.82 | 10.80 | 0.42 | | 458862.statgenpro | kelvin | large | GATK_YJ01 | 0 | 2016-10-28 11:09:05 | 49.90 | 10.30 | 0.24 | | 458860.statgenpro | kelvin | large | GATK_NG07 | 0 | 2016-10-28 11:19:56 | 53.44 | 10.70 | 0.40 | | 458858.statgenpro | kelvin | large | GATK_KC04 | 0 | 2016-10-28 10:57:39 | 55.96 | 10.30 | 0.11 | | 458857.statgenpro | kelvin | large | GATK_LR06 | 0 | 2016-10-28 10:54:35 | 55.46 | 10.30 | 0.11 | +-------------------+----------+-------+------------------------+------+---------------------+--------+-------+--------+ E S = Exit Status ; % = usage percentage; wtime = walltime
刪除工作(Delete a Job)
$ qdel JobID
例子:
[itsupport@omics]$qdel 1651
如何使用多個 CPU 核心加快檔案壓縮?
預設情況下,7z 進行檔案壓縮時通常只會使用單一 CPU 核心。當需要壓縮大量檔案或大型資料集時,整個過程可能相當耗時。
您可透過啟用多執行緒(Multi-threading)功能,讓 7z 同時使用多個 CPU 核心,以提升壓縮效能。
方法是指定 BZip2 壓縮演算法(-mm=BZip2),並使用 -mmt= 參數設定可使用的 CPU 執行緒數量。<#THREAD>” argument.
例如,以下指令會使用 4 個執行緒 進行壓縮:
7za a -mm=Bzip2 -mmt=4 output_zip_file input_file
如何檢查已使用的儲存空間?
用戶可使用 df 指令查看其主目錄(Home Directory)的磁碟配額及使用情況。例如:
[tmchan@omics ~]$df -h ~/ Filesystem Size Used Avail Use% Mounted on compellent2:/home 1000G 308G 693G 31% /home
如何使用 VS Code 連線至 HPCF?
由於 HPCF 採用特殊的安全存取設定,使用 VS Code Remote SSH 連線前,需要先在本機的 SSH 設定檔(SSH Config File)加入以下內容:
Host omics
HostName omics
User username
ProxyCommand ssh -q -W %h:%p username@omics.cpos.hku.hk
注意: 請將上述設定中的 username 替換為您的 HPCF2 用戶帳戶名稱。
您可參考下圖,在 VS Code 中直接開啟及編輯 SSH 設定檔(SSH Config File)。 常見位置如下:

為什麼 VS Code 無法連線至 HPCF?
由於 HPCF 的安全設定與 VS Code 最新版本的 Remote SSH 相容性問題,部分版本的 VS Code 可能無法正常連線至集群。
建議使用 VS Code 1.98.2 Portable 版本 連接 HPCF。由於 Portable 版本不會自動升級,可避免因版本更新導致連線失敗。
https://update.code.visualstudio.com/1.98.2/win32-x64-archive/stable
filename: VSCode-win32-x64-1.98.2.zip
md5sum: db4fefdae4986ef4ba5ba4524d8488cd
為避免 VS Code 升級後再次出現連線問題,建議停用自動更新功能。 請依照以下步驟設定: File > Preferences > Settings and search for “Update”, then:
- Uncheck “Update: Enable Windows Background Updates”
- Set “Update: Mode” to “none”

如何在 HPCF2 上啟動 RStudio / Jupyter Notebook?
如您的電腦尚未設定 X11 Forwarding,請先按照以下步驟進行設定:
- 安裝MobaXterm
- 前往官方網站下載安裝程式: https://mobaxterm.mobatek.net/
- 安裝完成後,啟動 MobaXterm
- 檢查 "X Server" 是否已啟動。

- 如顯示為停止狀態(Stopped),請按一下 "X Server" 啟動服務。

RStudio
- 登入 HPCF2(又稱 Omics)後,可使用以下指令查看系統內可用的 RStudio 軟體模組及版本:
ml av rstudio

- 執行以下指令載入 RStudio 軟體模組(建議使用 RStudio/1.4.1717 版本,以確保最佳相容性):RStudio/1.4.1717 for best compatibility):
ml RStudio/1.4.1717

當畫面顯示 “RStudio/1.4.1717 is loaded” ,表示 RStudio 模組已成功載入
註: 您亦可同時載入其他 R 模組(例如 R/4.2.1),以在 RStudio 中使用較新版本的 R。若未有額外載入 R 模組,系統預設使用 R 3.5.2。
- 輸入以下指令啟動 RStudio:
rstudio
執行後,終端機(Terminal)視窗可能會顯示以下訊息:

同時,RStudio 圖形介面視窗應會自動開啟:

Jupyter Notebook
- 輸入以下指令查看系統中可用的 Jupyter Notebook 模組及版本:
ml av ju

- 輸入以下指令載入 Jupyter Notebook:
ml jupyter/notebook

成功載入後,畫面應顯示以下訊息:
“firefox/latest is loaded
jupyter/notebook/ is loaded”.
- 輸入以下指令啟動 Jupyter Notebook:
jupyter notebook
執行後,終端機(Terminal)視窗可能會顯示以下訊息:

同時,系統應會開啟一個新的 Firefox 視窗,並載入 Jupyter Notebook 介面。

Jupyter Notebook 啟動後,請開啟一個新的 SSH 終端機視窗,並使用以下其中一個指令設定 SSH Local Port Forwarding:
ssh -N -f -L localhost:8889:omics:8889 username@hpcf2.cpos.hku.hk or ssh -N -f -L 8889:omics:8889 username@hpcf2.cpos.hku.hk or ssh -N -f -L 127.0.0.1:8889:omics:8889 username@hpcf2.cpos.hku.hk
8889 僅為範例連接埠(Port Number)。請根據 Jupyter Notebook 啟動時顯示的實際連接埠號碼進行修改。
如何解決連線 HPCF2 時出現的 SSH 主機識別警告(SSH Host Identification Warning)?
登入 HPCF2 時,您可能會遇到以下訊息:

此訊息表示 HPCF 伺服器的 SSH 主機金鑰(Host Key)已變更。一般而言,這是因伺服器完成安全維護或系統更新後重新產生 SSH 金鑰所致。
注意: 如果您自 2024 年 10 月起未曾登入 HPCF,出現此訊息屬預期情況,並非安全問題。
解決方法:
您需要先從本機電腦移除舊的 SSH 主機金鑰記錄。
請開啟終端機(Terminal)並執行以下指令:
ssh-keygen -R hpcf2.cpos.hku.hk

完成後,重新登入 HPCF2:
ssh your_username@hpcf2.cpos.hku.hk
(請將 your_username 替換為您的實際用戶名稱。)
當系統顯示以下提示訊息: “Are you sure you want to continue connecting (yes/no/[fingerprint])?” yes 請輸入 yes,然後按下 Enter 鍵。之後您便可輸入密碼並登入 HPCF2。

聯絡我們
itsupport.cpos@hku.hk
地址
香港薄扶林
沙宣道 5 號
香港賽馬會跨學科研究大樓 6 樓
電話:2831-5500
傳真:2818-5653
網址:https://cpos.hku.hk
電郵: enquiry.cpos@hku.hk
辦公時間
星期一至五:上午九時至下午五時半
當中下午一時至兩時 - 不設樣本及貨品接收
星期六、日、大學假期和公眾假期休息

